用数学方法检测大模型推理中的固执偏见,判断它是否真在追求真相。
Martingale Score: An Unsupervised Metric for Bayesian Rationality in LLM Reasoning
- 基于贝叶斯统计的鞅性质,设计无监督评分来衡量信念更新是否理性
- 发现多数模型在开放问题中存在信念固化,当前观点能预判未来改变
- 该评分可预测真实答案正确率,适合评估无法获取标准答案的推理过程
近期推理技术显著提升了大语言模型(LLMs)的表现,但也引发对其是否真正追求真理的担忧。本文提出一种基于贝叶斯统计中鞅性质的系统性评估框架,用于检测模型推理中的信念固化现象。该性质表明:在理性信念更新下,未来信念的期望值应等于当前信念,即更新不可从当前信念预测。我们提出无监督、基于回归的鞅分(Martingale Score),量化对这一性质的违背程度,反映模型对新证据的理性更新能力。在事件预测、价值判断类问题和学术论文评审等开放领域,我们发现信念固化现象普遍存在——当前信念能正向预测未来的信念变化。进一步识别出更易出现此现象的模型、推理方法与任务类型。最后通过实证验证,该分数在有真实标签的领域能有效预测准确性,说明即使在无真实答案场景下,鞅分也可作为推理过程可信度的有效代理指标。
原文摘要 · Abstract (English)
Recent advances in reasoning techniques have substantially improved the performance of large language models (LLMs), raising expectations for their ability to provide accurate, truthful, and reliable information. However, emerging evidence suggests that iterative reasoning may foster belief entrenchment and confirmation bias, rather than enhancing truth-seeking behavior. In this study, we propose a systematic evaluation framework for belief entrenchment in LLM reasoning by leveraging the Martingale property from Bayesian statistics. This property implies that, under rational belief updating, the expected value of future beliefs should remain equal to the current belief, i.e., belief updates are unpredictable from the current belief. We propose the unsupervised, regression-based Martingale Score to measure violations of this property, which signal deviation from the Bayesian ability of updating on new evidence. In open-ended problem domains including event forecasting, value-laden questions, and academic paper review, we find such violations to be widespread across models and setups, where the current belief positively predicts future belief updates, a phenomenon which we term belief entrenchment. We identify the models, reasoning techniques, and domains more prone to belief entrenchment. Finally, we validate the Martingale Score by showing that it predicts ground-truth accuracy on problem domains where ground truth labels are available. This indicates that, while designed as an unsupervised metric that operates even in domains without access to ground truth, the Martingale Score is a useful proxy of the truth-seeking ability of a reasoning process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。