提升大模型推理过程的可信度,让中间步骤更准确。
Trustworthy Reasoning: Evaluating and Enhancing Factual Accuracy in LLM Intermediate Thought Processes
- 用反事实数据训练事实核查器,识别推理链中的细节错误。
- 改进强化学习算法,使模型在事实性、连贯性上全面提升。
- 通过神经激活分析,揭示如何优化模型推理路径。
我们提出一个新框架,解决大语言模型(LLMs)在得出正确最终答案时,中间推理步骤仍存在事实性错误这一关键漏洞。该现象在医疗、法律和科研等高风险领域极具隐患,因错误但自信的推理可能误导用户做出危险决策。框架包含三部分:(1) 基于反事实增强数据训练的事实核查分类器,用于检测推理链中细微事实不一致;(2) 改进的组相对策略优化(GRPO)强化学习方法,通过多维奖励平衡事实性、连贯性和结构正确性;(3) 机制可解释方法,分析事实性提升在模型激活中的表现。跨多个前沿模型的评估显示,即使领先模型如Claude-3.7和GPT-o1,推理事实准确率也仅分别为81.93%和82.57%。本方法显著提升事实鲁棒性(最高改善49.90%),同时保持或优于Math-500、AIME-2024、GPQA等挑战性基准的表现。神经激活层面分析进一步揭示了事实增强如何重塑模型内部推理轨迹,为未来基于激活引导的训练方法奠定基础。
原文摘要 · Abstract (English)
We present a novel framework addressing a critical vulnerability in Large Language Models (LLMs): the prevalence of factual inaccuracies within intermediate reasoning steps despite correct final answers. This phenomenon poses substantial risks in high-stakes domains including healthcare, legal analysis, and scientific research, where erroneous yet confidently presented reasoning can mislead users into dangerous decisions. Our framework integrates three core components: (1) a specialized fact-checking classifier trained on counterfactually augmented data to detect subtle factual inconsistencies within reasoning chains; (2) an enhanced Group Relative Policy Optimization (GRPO) reinforcement learning approach that balances factuality, coherence, and structural correctness through multi-dimensional rewards; and (3) a mechanistic interpretability method examining how factuality improvements manifest in model activations during reasoning processes. Extensive evaluation across multi state-of-the-art models reveals concerning patterns: even leading models like Claude-3.7 and GPT-o1 demonstrate reasoning factual accuracy of only 81.93% and 82.57% respectively. Our approach significantly enhances factual robustness (up to 49.90% improvement) while maintaining or improving performance on challenging benchmarks including Math-500, AIME-2024, and GPQA. Furthermore, our neural activation-level analysis provides actionable insights into how factual enhancements reshape reasoning trajectories within model architectures, establishing foundations for future training methodologies that explicitly target factual robustness through activation-guided optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。