用贝叶斯追踪分离评分与排序,提升大模型推理可靠性
Prefix-Safe Bayesian Belief Tracking for LLM Reasoning Reliability:Separating Calibration from Ranking

- 基于前缀安全观测,递归更新双状态信念进行可靠性估计
- 结构感知信号在硬数学题上使AUROC提升0.110,优于纯评分方法
- 适合关注推理过程可信度、需在线校准的AI系统开发者
长推理链条需在最终答案生成前获得可靠性评估。本文研究前缀条件下的最终成功概率估计 $P(y=1 ackslashmid o_{1:t})$,采用前缀安全观测。序列贝叶斯信念追踪(SBBT)校准观测似然并递归更新双状态信念,可统一处理标量得分、文本标记、自验证信号、隐藏聚类、分词池探针及隐轨迹特征。在MATH-500、GSM8K、AIME 2025和RIMO-N的开源生成推理链上,概率质量与排序性能分离:仅使用得分的SBBT常改善Brier分数,而AUROC提升需依赖超越强前缀安全基线的结构感知证据。在最强硬数学场景中,结构感知观测使AUROC相比标准前缀安全基线提升+0.110。同前缀分类器审计下,MATH-500文本标记与RIMO-N自验证信号仍为正向。结果支持SBBT作为校准感知的在线推理框架,并揭示证据范式:标量得分主要支撑概率质量,而结构感知前缀信号仅在强前缀安全基线未吸收排名信息时才提升排序能力。
原文摘要 · Abstract (English)
Long reasoning traces need reliability estimates before final answers are known. We study prefix-conditioned eventual-success estimation, $P(y=1 \mid o_{1:t})$, using prefix-safe observations. Sequential Bayesian Belief Tracking (SBBT) calibrates observation likelihoods and recursively updates a two-state belief, providing a common tracker for scalar scores, text and self-verification markers, hidden clusters, token-pooling probes, and latent-trajectory features. Across generated open-weight traces on MATH-500, GSM8K, AIME 2025, and RIMO-N, probability quality and ranking separate: score-only SBBT often improves Brier, while AUROC gains require structure-aware evidence beyond strong prefix-safe baselines. In the strongest hard math setting, structure-aware observations reach +0.110 AUROC against standard prefix-safe baselines. Under a same-prefix classifier audit, MATH-500 text markers and RIMO-N self-verification signals remain positive. Together, these findings support SBBT as a calibration-aware online inference framework and expose an evidence regime: scalar scores mainly support probability quality, while structure-aware prefix signals support ranking only when strong prefix-safe baselines have not already absorbed the rank evidence.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。