构建首个聚焦推理的多轮对话检索基准,提升真实场景信息获取能力
RECOR: Reasoning-focused Multi-turn Conversational Retrieval Benchmark
- 通过分解与验证框架生成基于事实的多轮对话,确保每轮推理可追溯
- 结合对话历史与推理使检索性能提升一倍(nDCG@10从0.236升至0.479)
- 适合研究对话式搜索、逻辑推理与信息检索交叉方向的学者
现有基准将多轮对话与推理型检索分开评估,但现实信息获取需二者结合。为此,我们构建了一个面向推理型对话式信息检索的基准,包含跨十一领域的707个对话(共2,971轮)。为保证质量,采用分解与验证框架,将复杂查询转化为基于事实的多轮对话,通过多层级验证确保原子事实与源文档一致,并为每轮生成显式检索推理过程。全面评估显示,结合对话历史与推理使检索性能翻倍(基线nDCG@10 0.236 → 历史+推理 0.479),且专用推理模型显著优于密集编码器。然而分析表明,当逻辑关系未明确表达时,隐式推理仍具挑战性。
原文摘要 · Abstract (English)
Existing benchmarks treat multi-turn conversation and reasoning-intensive retrieval separately, yet real-world information seeking requires both. To bridge this gap, we present a benchmark for reasoning-based conversational information retrieval comprising 707 conversations (2,971 turns) across eleven domains. To ensure quality, our Decomposition-and-Verification framework transforms complex queries into fact-grounded multi-turn dialogues through multi-level validation, where atomic facts are verified against sources and explicit retrieval reasoning is generated for each turn. Comprehensive evaluation reveals that combining conversation history with reasoning doubles retrieval performance (Baseline .236 $\rightarrow$ History+Reasoning .479 nDCG@10), while reasoning-specialized models substantially outperform dense encoders. Despite these gains, further analysis highlights that implicit reasoning remains challenging, particularly when logical connections are not explicitly stated in the text.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。