现有自信度指标只看表面流畅,不关心逻辑结构。
Reasoning or Fluency? Dissecting Probabilistic Confidence in Best-of-N Selection
- 用三类干扰破坏推理步骤间的因果联系,保持语言流畅。
- 即使切断前后步骤关联,选择准确率仍基本不变。
- 提出新指标,专门捕捉推理步骤间的因果依赖关系。
概率自信度指标被广泛用于最佳- N 选择中作为推理质量的代理,其假设是更高的自信度意味着更可靠的推理过程。本文挑战这一假设,探究这些指标是否真正捕捉到推理步骤间的因果依赖关系。我们设计了三类系统性干扰,破坏推理步骤之间的因果联系,同时保持局部语言流畅性。在多种模型家族和推理基准上,发现即便引入严重干扰(如硬注意力掩码阻止模型关注前序步骤),选择准确率也仅轻微下降。结果表明,当前的概率指标对逻辑结构不敏感,主要反映表面流畅性或分布内先验。为此,我们提出一种对比因果度量,显式分离步骤间的因果依赖,并证明其在输出选择上比现有基于概率的方法更具可信度。
原文摘要 · Abstract (English)
Probabilistic confidence metrics are increasingly adopted as proxies for reasoning quality in Best-of-N selection, under the assumption that higher confidence reflects higher reasoning fidelity. In this work, we challenge this assumption by investigating whether these metrics truly capture inter-step causal dependencies necessary for valid reasoning. We introduce three classes of inter-step causality perturbations that systematically disrupt dependencies between reasoning steps while preserving local fluency. Surprisingly, across diverse model families and reasoning benchmarks, we find that selection accuracy degrades only marginally under these disruptions. Even severe interventions, such as applying hard attention masks that directly prevent the model from attending to prior reasoning steps, do not substantially reduce selection performance. These findings provide strong evidence that current probabilistic metrics are largely insensitive to logical structure, and primarily capture surface-level fluency or in-distribution priors instead. Motivated by this gap, we propose a contrastive causality metric that explicitly isolates inter-step causal dependencies, and demonstrate that it yields more faithful output selection than existing probability-based approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。