用共识推理图提升大模型思维链质量,让错误答案变正确。
Correct Prediction, Wrong Steps? Consensus Reasoning Knowledge Graph for Robust Chain-of-Thought Synthesis

- 构建多路径推理的共识知识图,自动过滤错误步骤
- 平均提升标签预测准确率超10%,逻辑与数学任务全胜
- 适合需要可靠推理过程的AI系统开发者
大模型的推理轨迹存在复杂缺陷——包括步骤内错误(如逻辑谬误、幻觉)和步骤层面错误(如过度思考、思考不足),且样本间差异显著。直觉上可通过提供真实标签来引导模型推理,但实验表明这无助于提升推理能力。为此,本文提出CRAFT框架,通过整合多个候选推理路径的共识部分构建推理知识图(RKG),并基于拓扑结构生成高质量推理链。该方法在多个逻辑与数学推理基准上平均提升标签预测准确率超过10%,全面超越所有基线模型。详细评估进一步验证其在推理轨迹质量多个维度上的持续改进。
原文摘要 · Abstract (English)
LLM reasoning traces suffer from complex flaws -- *Step Internal Flaws* (logical errors, hallucinations, etc.) and *Step-wise Flaws* (overthinking, underthinking), which vary by sample. A natural approach would be to provide ground-truth labels to guide LLMs' reasoning. Contrary to intuition, we show that this yields no improvement in reasoning ability. We then propose CRAFT, a unified framework that mitigates both types of Step flaws, which builds a Reasoning Knowledge Graph (RKG) based on the consensus parts of multiple candidate traces, and synthesizes a high-quality trace through topological generation. Our approach improves label-prediction accuracy by 10+% on average, and consistently outperforms all baselines across both logical and mathematical reasoning benchmarks. Further, detailed benchmark evaluation proves that our method also improves the quality of LLMs' reasoning traces in multiple dimensions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。