大模型推理出错常因虚构题目关键信息,而非逻辑错误。
Reasoning Large Language Model Errors Arise from Hallucinating Critical Problem Features
- 通过分析推理过程文本,发现模型会凭空编造题目未给的图边。
- 在不同复杂度下,多款模型因虚构信息导致错误占比超一半。
- 适合关注模型可靠性与提示工程的研究者和开发者参考。
大型语言模型在链式思维(CoT)策略与强化学习训练下,推理能力显著提升;然而这些推理型大模型(RLLMs)仍存在缺陷,理解其失败模式对用户和开发者至关重要。我们在图着色这一可变复杂度的约束满足逻辑问题上,测试了o1-mini、o3-mini、DeepSeek-R1、Claude 3.7 Sonnet、Gemini 2.5 Pro Preview和Grok 3 Mini Beta,发现错误率对比与链式思维/解释文本分析均表明,RLLMs容易在提示中未指定的情况下幻觉生成图边。该现象存在于多个问题复杂度层级和语义框架中,并显著导致各模型的错误答案,对部分模型而言甚至占绝大多数。我们还在稳定匹配问题的小规模实验中验证了此输入冲突性幻觉的泛化能力。结果表明,RLLMs可能普遍存在对问题细节误表征的问题,并提出了若干设计改进建议以缓解该缺陷。
原文摘要 · Abstract (English)
Large language models have recently made great strides in reasoning task performance through chain-of-thought (CoT) strategies trained via reinforcement learning; however, these "reasoning large language models" (RLLMs) remain imperfect reasoners, and understanding the frequencies and causes of their failure modes is important for both users and developers. We test o1-mini, o3-mini, DeepSeek-R1, Claude 3.7 Sonnet, Gemini 2.5 Pro Preview, and Grok 3 Mini Beta on graph coloring as a variable-complexity constraint-satisfaction logic problem, and find evidence from both error rate comparisons and CoT/explanation text analysis that RLLMs are prone to hallucinate graph edges not specified in the prompt. This phenomenon persists across multiple problem complexity levels and semantic frames, and it appears to account for a significant fraction of the incorrect answers from every tested model, and the vast majority of them for some models. We also validate the generalizability of this input-conflicting hallucination phenomenon with smaller-scale experiments on a type of stable matching problem. Our results indicate that RLLMs may possess broader issues with misrepresentation of problem specifics, and we offer suggestions for design choices to mitigate this weakness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。