source: arxiv machine learning: mirror horizon: viable path entropy as a measure of bounded reflection

level: research

mirror theory suggests that an intelligent system's capabilities are not just about what it represents, but about the coherent continuations it can sustain when reflecting on its own outputs. this paper introduces viable path entropy (vpe), a measure that quantifies this capacity. vpe works by giving a system a mirror state, a rollout protocol, a verifier, and a mode map. it then measures two things: the chance of reaching a viable continuation, and the variety of verified continuation modes among successful rollouts. this breaks down bounded capability into probability and diversity components.

the theoretical framework behind vpe includes several key ideas. intuition acts as a local constraint that leaves some options undetermined. taste serves as a pressure that selects invariants. reflection is the process of resolving underdetermination guided by taste. geometry is the learned structure that makes future reflection stable. these concepts together form a scaffold for understanding how systems can maintain coherent self-reflection over multiple steps.

the authors tested this theory using language model reasoning experiments on the gsm8k dataset. the experiments show how vpe can be applied to measure a model's ability to generate diverse, verified reasoning paths. by focusing on verified continuations, vpe provides a practical way to assess not just accuracy, but the robustness and flexibility of a model's reasoning under self-reflection. this moves beyond simple output evaluation to consider the stability of the reasoning process itself.

why it matters: vpe offers a new way to evaluate ai reasoning systems by measuring their capacity for stable self-reflection, which is crucial for building reliable models that can check and refine their own outputs.


source: arxiv machine learning: mirror horizon: viable path entropy as a measure of bounded reflection