arXiv:2608.03527cs.IRcs.AI2026-08

用搜索标准指导文档重排,让检索结果更符合研究需求

Training Documents Reranker with Search Rubrics for Deep Research Agent

论文配图:Training Documents Reranker with Search Rubrics for Deep Research Agent
图 1 · 摘自论文原文
  • 基于大模型构建分层搜索标准,明确优质文档集应满足的条件
  • 在四个深度研究基准上比最强基线提升2.6分,跨任务泛化能力强
  • 适合需要精准、全面、权威信息的研究型AI系统使用

检索系统通过提供相关文档帮助深度研究代理生成高质量答案。然而,现有检索器通常仅通过相关性匹配选择文档,单独匹配度高的前k篇文档组合起来未必能满足代理查询的复杂信息需求(如多样性、简洁性与权威性)。本文提出面向搜索的评估标准,明确每类查询下高质量文档集合应满足的要求。这些标准采用分层结构,并由强大语言模型生成。基于此,我们进一步训练了一个文档重排器RubricRanker,从检索结果中选出高质量子集。设计了两阶段训练框架:标准引导的监督微调与基于标准的强化学习。大量实验表明,RubricRanker在四个深度研究基准上优于最强基线2.6分,并在五个RAG基准上展现出良好泛化能力。

原文摘要 · Abstract (English)

Retrieval systems help deep research agents generate high-quality answers by providing relevant documents. However, existing retrievers typically select documents through relevance matching, while individually well-matched top-$k$ documents may not form a \textit{set} that satisfies the complex information needs of an agent query (\eg, diverse, concise and authoritative documents). In this paper, we propose search-oriented rubrics that \textit{explicitly} define the requirements that high-quality document sets should satisfy for each agent query. Our search rubrics are organized into a hierarchical structure and synthesized using a powerful LLM. Based on these search rubrics, we further train a document reranker \textbf{RubricRanker} to select a high-quality subset from retrieved documents. We design a two-stage training framework that consists of rubrics-guided supervised fine-tuning and rubric-based reinforcement learning. Extensive experiments demonstrate that RubricRanker outperforms the strongest baseline by 2.6 points on four deep research benchmarks and generalizes well to five RAG benchmarks.

检索优化大模型应用研究代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。