提出时空双视角框架GeoFaith,解决大模型推理不忠实问题。
GeoFaith: A Spatio-Temporal Dual View of Faithful Chain-of-Thought

- 利用隐空间几何结构与熵动态诊断推理过程是否忠实
- 扩展标注样本至2万条,训练出超越GPT-5的80亿参数检测器
- 支持可解释性更强、更短且准确率不降的推理链生成
链式思维(CoT)推理推动了大语言模型的发展,但基于结果的监督导致普遍出现事后合理化,产生看似合理却不可信的推理链。现有忠实度评估方法或难以扩展、成本高,或不可靠。我们提出GeoFaith,一种时空双视角框架,通过挖掘隐变量的几何结构与熵动态来诊断并强化推理忠实性。构建可扩展的自举流程,将步骤级标注从1000扩展至20000样本,覆盖四个领域;训练出一个80亿参数的忠实度检测器,在标准基准上表现优于GPT-5;设计一种兼顾结果正确性、过程忠实性与轨迹一致性的忠实感知强化学习框架。实验表明,该方法在忠实度检测与下游推理任务中均表现优异,生成更短、更易解释的推理链,且不牺牲准确性。代码将公开。
原文摘要 · Abstract (English)
Chain-of-Thought (CoT) reasoning has advanced large language models (LLMs), but outcome-based supervision leads to pervasive post-hoc rationalization, producing plausible yet unfaithful reasoning chains. Most prior faithfulness assessment methods are either unscalable, expensive, or unreliable. We propose GeoFaith, a spatio-temporal framework that leverages latent geometric structure and entropy dynamics to diagnose and enforce faithful reasoning. We develop a scalable bootstrapping pipeline expanding step-level annotations from 1k to 20k samples across four domains, train an 8B faithfulness detector outperforming GPT-5 on standard benchmarks, and design a faithfulness-aware reinforcement learning framework jointly optimizing outcome correctness, process faithfulness, and trajectory consistency. Experiments show the proposed method achieves superior performance on both faithfulness detection and downstream reasoning, producing shorter, more interpretable chains without sacrificing accuracy. Our code will be made available publicly.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。