arXiv:2508.04699cs.CLcs.AI2025-08被引 5

剖析推理模型在多跳问答中的失败原因,揭示其幻觉根源。

Hop, Skip, and Overthink: Diagnosing Why Reasoning Models Fumble during Multi-Hop Analysis

  • 从跳数、覆盖度、过度思考三维度分类错误模式。
  • 发现模型在多源信息整合时易遗漏关键内容或冗余推演。
  • 适合研究推理机制与提升AI可信度的学者与工程师。

推理模型在数学求解、深度搜索和抽取式问答等复杂任务中取得突破,但其为何比通用语言模型更易产生幻觉尚不明确。本文系统研究当前语言模型在多跳问答任务中的推理失败问题,提出一种新颖的细粒度错误分类框架,从三个关键维度分析:涉及文档的多样性与独特性(“跳”)、相关信息捕获的完整性(“覆盖度”)以及认知效率低下(“过度思考”)。通过人工标注结合自动化指标的严格评估,揭示了传统精度评价掩盖的复杂错误模式。该研究深化了对现有模型认知局限的理解,为未来提升推理一致性、透明度与鲁棒性提供了可操作指引。

原文摘要 · Abstract (English)

The emergence of reasoning models and their integration into practical AI chat bots has led to breakthroughs in solving advanced math, deep search, and extractive question answering problems that requires a complex and multi-step thought process. Yet, a complete understanding of why these models hallucinate more than general purpose language models is missing. In this investigative study, we systematicallyexplore reasoning failures of contemporary language models on multi-hop question answering tasks. We introduce a novel, nuanced error categorization framework that examines failures across three critical dimensions: the diversity and uniqueness of source documents involved ("hops"), completeness in capturing relevant information ("coverage"), and cognitive inefficiency ("overthinking"). Through rigorous hu-man annotation, supported by complementary automated metrics, our exploration uncovers intricate error patterns often hidden by accuracy-centric evaluations. This investigative approach provides deeper insights into the cognitive limitations of current models and offers actionable guidance toward enhancing reasoning fidelity, transparency, and robustness in future language modeling efforts.

推理模型多跳问答错误分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。