arXiv:2509.25941cs.AIcs.CL2025-09

通过评估题目可解性,提升大模型推理的准确性与可靠性。

Boosting Process-Correct CoT Reasoning by Modeling Solvability of Multiple-Choice QA

  • 基于题目可解性动态调整训练目标,避免无效推理链
  • 在数学和多模态数据集上,过程正确率显著提升
  • 适合关注推理可信度与幻觉抑制的研究者

大语言模型的推理质量不仅取决于答案正确性,还依赖于中间推理步骤的有效性。我们通过多项选择题问答(MCQA)研究这一问题,该任务提供固定选项的可控环境。分析发现,当题目对模型而言实际不可解时,更易产生虚假的思维链(CoT),导致假阳性结果。通过估计每道题的可解性,我们发现存在一个中间区间,此时学习效果最佳。基于此,我们改进了基于结果监督的奖励模型和群体相对优势的强化学习方法,将可解性纳入优化目标。在数学和多模态数据集上的实验表明,这些改进稳定提升了过程正确推理的比例,且在强化学习中也提高了最终答案准确率。结果表明,可解性是减少幻觉、增强思维链推理可靠性的关键因素。

原文摘要 · Abstract (English)

Reasoning quality in large language models depends not only on producing correct answers but also on generating valid intermediate steps. We study this through multiple-choice question answering (MCQA), which provides a controlled setting with fixed answer options. Our analysis shows that when questions are effectively unsolvable for a model, spurious chains of thought (CoTs) are more likely to appear, leading to false positives. By estimating the solvability of each question, we uncover an intermediate regime where learning is most effective. Building on this insight, we adapt outcome-supervised reward models and reinforcement learning with group-relative advantage to incorporate solvability into their objectives. Across experiments on math and multimodal datasets, these modifications consistently yield higher rates of process-correct reasoning and, in reinforcement learning, improved answer accuracy as well. Our results highlight solvability as a key factor for reducing hallucinations and increasing reliability in CoT reasoning.

推理质量思维链可解性建模强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。