研究大模型推理中错误信息如何传播及修复难题
Unraveling Misinformation Propagation in LLM Reasoning
- 分析错误输入如何影响大模型数学推理的中间步骤
- 即使被告知纠正,模型纠错成功率不足50%,准确率下降10%至72%
- 早期介入纠正可显著减少错误传播,微调数据效果更佳
大型语言模型在推理方面表现出色,被视为辅助人类解决问题的有力工具。然而,当用户因疏忽或知识盲区引入错误信息时,模型性能会受到怎样的影响?这类错误在真实交互中普遍存在,但其在模型推理过程中的传播机制尚未被充分研究。本文聚焦数学推理,全面分析错误信息对中间推理步骤和最终答案的影响,并考察模型在明确指令下纠正错误的能力。结果显示,即便拥有正确内部知识,模型在纠正错误时成功率仍低于一半,导致准确率下降10.02%至72.20%,即使使用思维链模型也存在4.30%至19.97%的性能退化。进一步分析表明,尽早进行事实修正能最有效抑制错误传播;通过合成数据进行早期修正微调,显著提升了推理的准确性。本研究为缓解错误信息传播提供了实用策略。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated impressive capabilities in reasoning, positioning them as promising tools for supporting human problem-solving. However, what happens when their performance is affected by misinformation, i.e., incorrect inputs introduced by users due to oversights or gaps in knowledge? Such misinformation is prevalent in real-world interactions with LLMs, yet how it propagates within LLMs' reasoning process remains underexplored. Focusing on mathematical reasoning, we present a comprehensive analysis of how misinformation affects intermediate reasoning steps and final answers. We also examine how effectively LLMs can correct misinformation when explicitly instructed to do so. Even with explicit instructions, LLMs succeed less than half the time in rectifying misinformation, despite possessing correct internal knowledge, leading to significant accuracy drops (10.02% - 72.20%), and the degradation holds with thinking models (4.30% - 19.97%). Further analysis shows that applying factual corrections early in the reasoning process most effectively reduces misinformation propagation, and fine-tuning on synthesized data with early-stage corrections significantly improves reasoning factuality. Our work offers a practical approach to mitigating misinformation propagation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。