arXiv:2506.00264cs.CL2025-06ACL被引 7

构建多跳虚假前提问答数据集,评估大模型在复杂推理中的假设识别能力。

MultiHoax: A Dataset of Multi-hop False-Premise Questions

  • 设计多跳推理任务,需跨步骤验证前提一致性。
  • 主流大模型在多国、多领域任务中检测假前提准确率普遍偏低。
  • 适合关注模型可靠性与批判性推理能力的研究者使用。

随着大型语言模型在高风险领域的广泛应用,其识别错误假设和进行批判性推理的能力对确保输出可靠性至关重要。虚假前提问题(FPQ)是评估模型在错误假设导致错误回答时表现的重要方法。现有基准主要聚焦于单跳FPQ,但现实推理常需多跳推断,模型需验证多个推理步骤的一致性,而非依赖表面线索。为此,我们提出MultiHoax,一个用于评估大模型处理复杂多步推理中虚假前提能力的基准数据集。该数据集覆盖七个国家、十大知识类别,以维基百科为主要知识源,支持跨区域事实推理。实验表明,当前最先进大模型在不同国家、知识类别及多跳推理类型中均难以有效检测虚假前提,凸显了提升假前提识别能力和增强多跳推理鲁棒性的迫切需求。

原文摘要 · Abstract (English)

As Large Language Models are increasingly deployed in high-stakes domains, their ability to detect false assumptions and reason critically is crucial for ensuring reliable outputs. False-premise questions (FPQs) serve as an important evaluation method by exposing cases where flawed assumptions lead to incorrect responses. While existing benchmarks focus on single-hop FPQs, real-world reasoning often requires multi-hop inference, where models must verify consistency across multiple reasoning steps rather than relying on surface-level cues. To address this gap, we introduce MultiHoax, a benchmark for evaluating LLMs' ability to handle false premises in complex, multi-step reasoning tasks. Our dataset spans seven countries and ten diverse knowledge categories, using Wikipedia as the primary knowledge source to enable factual reasoning across regions. Experiments reveal that state-of-the-art LLMs struggle to detect false premises across different countries, knowledge categories, and multi-hop reasoning types, highlighting the need for improved false premise detection and more robust multi-hop reasoning capabilities in LLMs.

虚假前提多跳推理大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。