arXiv:2512.01725cs.CL2025-12AAAI被引 2

LLM在多解任务中因过度自信而遗漏答案,需用长链思维改进。

Beware of Reasoning Overconfidence: Pitfalls in the Reasoning Process for Multi-solution Tasks

  • 用长链思维逐步探索,避免过早锁定单一路径。
  • 短链思维下模型对不完整答案信心过强,错误率超40%。
  • 适合研究推理过程缺陷或评估模型全面性的人参考。

大型语言模型在单答案推理任务中表现优异,但在需要生成全面且多样答案的多解任务中表现不佳。我们将其归因于‘推理过度自信’:即在答案集不完整时仍表现出过度确定。为此,我们引入了多解问题基准测试集MuSoBench。实验表明,传统短链思维(Short-CoT)提示范式存在明显过度自信现象,而新兴的长链思维(Long-CoT)通过迭代探索与自我反思有效缓解该问题。我们进一步分析了可观察行为及影响因素。为探究根本原因,提出‘认知刚性假说’,认为过度自信源于推理过程过早收敛到有限思考路径。注意力熵分析为此提供了初步支持。这些发现为评估模型推理完整性提供了工具,并强调应将评测重点从单答案准确率转向全面探索能力。

原文摘要 · Abstract (English)

Large Language Models (LLMs) excel in reasoning tasks requiring a single correct answer, but they perform poorly in multi-solution tasks that require generating comprehensive and diverse answers. We attribute this limitation to \textbf{reasoning overconfidence}: a tendency to express undue certainty in an incomplete solution set. To examine the effect, we introduce \textit{MuSoBench}, a benchmark of multi-solution problems. Experiments show that the conventional short chain-of-thought (Short-CoT) prompting paradigm exhibits pronounced overconfidence, whereas the emerging long chain-of-thought (Long-CoT) approach mitigates it through iterative exploration and self-reflection. We further characterise observable behaviours and influential factors. To probe the underlying cause, we propose the \textbf{cognitive-rigidity hypothesis}, which posits that overconfidence arises when the reasoning process prematurely converges on a narrow set of thought paths. An attention-entropy analysis offers preliminary support for this view. These findings provide tools for assessing the completeness of LLM reasoning and highlight the need to move evaluation beyond single-answer accuracy toward comprehensive exploration.

大模型推理多解任务认知刚性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。