arXiv:2605.04018cs.CLcs.IR2026-05ACL被引 1

新基准+合成数据提升检索器在智能搜索中的推理能力

Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems

论文配图:Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
图 1 · 摘自论文原文
  • 构建多维度标注的BRIGHT-Pro基准,支持迭代搜索评估
  • 用分解式合成数据训练出更强的RTriever-4B模型
  • 适合研究智能搜索、推理增强检索的开发者

推理密集型检索旨在发现支持下游推理的证据,而非仅匹配主题相关性。这对需要迭代搜索与融合的智能搜索系统至关重要。现有工作在评估和训练上均受限:如BRIGHT基准仅提供单一黄金证据集,且孤立评估检索器;合成数据常只优化单段落相关性,而非证据组合构建。本文提出BRIGHT-Pro,一个由专家标注的扩展基准,为每个查询添加多方面黄金证据,并在静态与智能搜索协议下评估检索器。同时构建了RTriever-Synth,一个基于视角分解的合成语料库,生成互补正例与条件化难负例,用于对Qwen3-Embedding-4B的RTriever-4B进行LoRA微调。跨词汇、通用及推理密集型检索器的实验表明,视角感知与智能搜索评估能揭示标准指标掩盖的行为,且RTriever-4B显著优于基线模型。

原文摘要 · Abstract (English)

Reasoning-intensive retrieval aims to surface evidence that supports downstream reasoning rather than merely matching topical similarity. This capability is increasingly important for agentic search systems, where retrievers must provide complementary evidence across iterative search and synthesis. However, existing work remains limited on both evaluation and training: benchmarks such as BRIGHT provide narrow gold sets and evaluate retrievers in isolation, while synthetic training corpora often optimize single-passage relevance rather than evidence portfolio construction. We introduce BRIGHT-Pro, an expert-annotated benchmark that expands each query with multi-aspect gold evidence and evaluates retrievers under both static and agentic search protocols. We further construct RTriever-Synth, an aspect-decomposed synthetic corpus that generates complementary positives and positive-conditioned hard negatives, and use it to LoRA fine-tune RTriever-4B from Qwen3-Embedding-4B. Experiments across lexical, general-purpose, and reasoning-intensive retrievers show that aspect-aware and agentic evaluation expose behaviors hidden by standard metrics, while RTriever-4B substantially improves over its base model.

检索增强智能搜索大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。