构建多跳推理与模糊性交织的评测基准,揭示大模型在复杂推理中的瓶颈。
MARCH: Evaluating the Intersection of Ambiguity Interpretation and Multi-hop Inference
- 设计多阶段模糊问题生成流程,确保每题有多个合理推理路径。
- 2209道题目中,顶尖模型准确率不足50%,凸显挑战性。
- 提出双阶段代理框架,分离模糊处理与证据推理,性能显著提升。
现实世界的多跳问答天然伴随模糊性,单个问题可引发多个需独立解决的推理路径。由于模糊性可能出现在任何阶段,模型必须在整条推理链中应对层叠不确定性。尽管此类现象在真实用户查询中普遍存在,现有评测主要聚焦单跳模糊性,对多步推理与层级模糊性的交互仍研究不足。本文提出MARCH基准,包含2,209道通过多大语言模型验证并经人工标注、一致性强的多跳模糊问题。实验表明,即使最先进模型在该基准上表现不佳,证实结合模糊性解析与多步推理是重大挑战。为此,我们提出CLARION,一种两阶段代理框架,显式分离模糊性规划与证据驱动推理,显著优于现有方法,为构建鲁棒推理系统提供新路径。
原文摘要 · Abstract (English)
Real-world multi-hop QA is naturally linked with ambiguity, where a single query can trigger multiple reasoning paths that require independent resolution. Since ambiguity can occur at any stage, models must navigate layered uncertainty throughout the entire reasoning chain. Despite its prevalence in real-world user queries, previous benchmarks have primarily focused on single-hop ambiguity, leaving the complex interaction between multi-step inference and layered ambiguity underexplored. In this paper, we introduce MARCH, a benchmark for their intersection, with 2,209 multi-hop ambiguous questions curated via multi-LLM verification and validated by human annotation with strong agreement. Our experiments reveal that even state-of-the-art models struggle with MARCH, confirming that combining ambiguity resolution with multi-step reasoning is a significant challenge. To address this, we propose CLARION, a two-stage agentic framework that explicitly decouples ambiguity planning from evidence-driven reasoning, significantly outperforms existing approaches, and paves the way for robust reasoning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。