提出四维框架解析多跳问答的检索推理过程,揭示不同方法的权衡与趋势。
Retrieval--Reasoning Processes for Multi-hop Question Answering: A Four-Axis Design Framework and Empirical Trends
- 构建四轴分析框架:执行计划、索引结构、下一步控制、停止条件
- 在HotpotQA等数据集上发现有效性、效率与证据忠实性间的普遍权衡
- 适合研究多跳问答系统设计与评估的研究者参考
多跳问答需要系统迭代式地检索证据并跨多步进行推理。尽管近期RAG和代理型方法表现优异,但其底层的检索-推理过程常被隐含处理,导致不同模型家族间的流程选择难以比较。本文以执行过程为分析单元,提出涵盖(A)整体执行计划、(B)索引结构、(C)下一步控制策略与触发机制、(D)停止/继续判定标准的四轴框架。基于该框架,我们映射了代表性多跳问答系统,并综合多个标准基准(如HotpotQA、2WikiMultiHopQA、MuSiQue)上的消融实验与趋势,揭示有效性、效率与证据忠实性之间的常见权衡。最后提出开放挑战,包括结构感知规划、可迁移控制策略以及分布偏移下的鲁棒停止机制。
原文摘要 · Abstract (English)
Multi-hop question answering (QA) requires systems to iteratively retrieve evidence and reason across multiple hops. While recent RAG and agentic methods report strong results, the underlying retrieval--reasoning \emph{process} is often left implicit, making procedural choices hard to compare across model families. This survey takes the execution procedure as the unit of analysis and introduces a four-axis framework covering (A) overall execution plan, (B) index structure, (C) next-step control (strategies and triggers), and (D) stop/continue criteria. Using this schema, we map representative multi-hop QA systems and synthesize reported ablations and tendencies on standard benchmarks (e.g., HotpotQA, 2WikiMultiHopQA, MuSiQue), highlighting recurring trade-offs among effectiveness, efficiency, and evidence faithfulness. We conclude with open challenges for retrieval--reasoning agents, including structure-aware planning, transferable control policies, and robust stopping under distribution shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。