通过中间步骤一致性检测,提升大模型数学推理的可靠性。
Enhancing Mathematical Reasoning in Large Language Models with Self-Consistency-Based Hallucination Detection
- 设计结构化自一致性框架,校验推理中间步骤
- 在定理证明等任务中准确率提升,输出更稳定
- 适合需要高可信数学推理的科研与教育场景
大型语言模型在数学推理方面表现出色,但在定理证明、符号运算和数值计算中仍易产生看似合理却错误的幻觉。现有基于自一致性(SC)的方法主要应用于最终答案选择,忽视了中间推理步骤的一致性。本文提出一种结构化自一致性框架,强制要求中间步骤与最终结果保持逻辑一致,有效减少逻辑矛盾与幻觉。我们在三个核心数学任务上评估该方法:定理证明、符号变换和数值计算。实验表明,该方法显著提升了证明有效性、符号推理准确率和数值稳定性,同时保持高效计算。进一步分析显示,结构化自一致性不仅提高解题准确率,还降低了模型输出的方差。结果表明,自一致性是提升大模型数学推理能力的可靠机制,有助于构建更可靠、可解释的AI数学系统。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated strong mathematical reasoning capabilities but remain susceptible to hallucinations producing plausible yet incorrect statements especially in theorem proving, symbolic manipulation, and numerical computation. While self-consistency (SC) has been explored as a means to improve factuality in LLMs, existing approaches primarily apply SC to final-answer selection, neglecting the logical consistency of intermediate reasoning steps. In this work, we introduce a structured self-consistency framework designed to enhance the reliability of mathematical reasoning. Our method enforces self-consistency across intermediate steps and final outputs, reducing logical inconsistencies and hallucinations. We evaluate our approach across three core mathematical tasks: theorem proving, symbolic transformation, and numerical computation. Experimental results demonstrate that SC significantly improves proof validity, symbolic reasoning accuracy, and numerical stability while maintaining computational efficiency. Further analysis reveals that structured self-consistency not only enhances problem-solving accuracy but also reduces the variance of model-generated outputs. These findings highlight self-consistency as a robust mechanism for improving mathematical reasoning in LLMs, paving the way for more reliable and interpretable AI-driven mathematics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。