提升大模型推理准确性,通过多轮自洽性评估纠正事实幻觉。
Reasoning Error from Known Fact: Step-Level Self-Consistency Group Relative Policy Optimization for LLM

- 基于多轮推理结果的步骤级自洽性评分,动态分配奖励。
- 在数学推理和幻觉检测榜单上均达当前最佳表现。
- 适合关注大模型可信推理与幻觉控制的研究者。
随着大语言模型(LLMs)的快速发展,现代系统不仅能具备强大的基础能力与广泛知识,还能通过长链多步推理解决复杂问题。然而,随着推理路径变长,模型在推理过程中可能产生大量幻觉内容,且难以被察觉。本文对大模型推理中的幻觉进行细粒度分析,发现其特别容易出现上下文敏感的事实幻觉:即模型实际上掌握相关知识,却因推理过程中的上下文干扰而犯错。为此,我们提出步骤级自洽性群体相对策略优化(SSC-GRPO),通过在多个推理轨迹中计算各步骤的自洽性得分,为推理路径分配步骤级奖励。相比先前方法,SSC-GRPO在数学推理基准与幻觉排行榜上均达到最先进性能。研究结果为检测与缓解大模型推理过程中的幻觉提供了新视角。
原文摘要 · Abstract (English)
With the rapid advancement of large language models (LLMs), modern systems not only possess strong foundational capabilities and extensive knowledge, but can also solve complex problems via long, multi-step reasoning. However, as reasoning traces become longer, LLMs may produce a substantial amount of hallucinated content during the reasoning process, which is often difficult to detect. In this work, we conduct a fine-grained analysis of hallucinations arising in LLM reasoning and find that the reasoning traces are particularly prone to Context-Sensitive Factual Hallucinations: cases where the model actually has the relevant knowledge, yet makes factual errors due to contextual interference during reasoning. To address this issue, we propose Step-level Self-Consistency Group Relative Policy Optimization (SSC-GRPO), which assigns step-level rewards to reasoning traces by computing self-consistency scores of individual steps across multiple rollouts. Compared with prior methods, SSC-GRPO achieves state-of-the-art performance on both mathematical reasoning benchmarks and hallucination leaderboards. Our results offer a new perspective for detecting and mitigating hallucinations in the reasoning process of large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。