arXiv:2510.24794cs.CL2025-10ACL被引 3

让大模型推理更可信,通过优化思维过程提升答案准确性。

MR-Align: Meta-Reasoning Informed Factuality Alignment for Large Reasoning Models

  • 用元推理分析思维链,动态调整每步思考的权重
  • 在4个问答数据集上准确率提升,错误推理减少30%以上
  • 无需外部验证器,适合追求高可信推理的应用场景

大型推理模型(LRMs)在复杂推理任务中表现强劲,但在依赖证据的事实性问题上提升有限。我们发现这主要源于‘推理-答案匹配差距’:模型在推理中识别出正确事实,却未能将其融入最终回答,降低事实一致性。为此,我们提出MR-ALIGN框架,一种基于元推理的对齐方法,无需外部验证器即可增强事实性。该方法量化模型思维过程中状态转移的概率,构建感知转移的隐式奖励信号,强化有益的推理模式,抑制缺陷性思维片段。通过重新加权,将词级信号转化为概率感知的段落评分,引导更连贯、更利于事实正确的推理路径。在四个事实性问答数据集和一个长文本事实性基准上的实证评估表明,MR-ALIGN在保持一致性的前提下持续提升准确率与真实性,并显著减少误导性推理。结果表明,对推理过程本身的对齐比仅对输出对齐更为关键,是推动LRMs事实性进步的核心。

原文摘要 · Abstract (English)

Large reasoning models (LRMs) show strong capabilities in complex reasoning, yet their marginal gains on evidence-dependent factual questions are limited. We find this limitation is partially attributable to a reasoning-answer hit gap, where the model identifies the correct facts during reasoning but fails to incorporate them into the final response, thereby reducing factual fidelity. To address this issue, we propose MR-ALIGN, a Meta-Reasoning informed alignment framework that enhances factuality without relying on external verifiers. MR-ALIGN quantifies state transition probabilities along the model's thinking process and constructs a transition-aware implicit reward that reinforces beneficial reasoning patterns while suppressing defective ones at the atomic thinking segments. This re-weighting reshapes token-level signals into probability-aware segment scores, encouraging coherent reasoning trajectories that are more conducive to factual correctness. Empirical evaluations across four factual QA datasets and one long-form factuality benchmark show that MR-ALIGN consistently improves accuracy and truthfulness while reducing misleading reasoning. These results highlight that aligning the reasoning process itself, rather than merely the outputs, is pivotal for advancing factuality in LRMs.

推理对齐事实性元推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。