让检索结果更优:用评分标准直接指导文档集选择与排序
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking

- 基于评分标准构建文档集评估框架,捕捉文档间冗余、互补等关系
- 12种重排器最高仅覆盖45%关键维度,跨文档协作能力普遍薄弱
- 无需训练的新方法,用评分标准生成选文信号,少查几轮也能更好
随着大语言模型和AI代理成为搜索结果的主要使用者,文档集质量决定了下游生成的上限。现有评估系统仍局限于独立打分并用nDCG聚合,忽略了文档间的相互作用(冗余、冲突、互补),无法回答为何一个文档集优于另一个。为此,我们提出一个完整的评估-诊断-优化框架。设计SetwiseEvalKit,一个涵盖短文本与长文本场景的三级九维评估基准,包含约28,000条高质量评估评分标准。系统性评估12个重排器:即使最优方法覆盖率也不超过45%,跨文档协调维度普遍表现不佳,且无单一方法在两种场景下均保持领先。在此基础上,提出Rubric4Setwise——一种无需训练的方法,将评分标准转化为文档集选择信号,在更少文档和搜索轮次下实现最佳下游生成效果。它是唯一在两种场景下均保持顶尖表现的方法,验证了从评估到优化闭环的有效性。
原文摘要 · Abstract (English)
As large language models and AI agents become the primary consumers of search results, document set quality determines the upper bound of downstream generation. Yet existing evaluation systems remain confined to scoring documents independently and aggregating via nDCG, ignoring inter-document interactions (redundancy, conflict, complementarity) and unable to answer what makes one document set better than another. To address these issues, we propose a complete evaluate-diagnose-optimize framework. We design SetwiseEvalKit, a three-level, nine-dimension document set evaluation benchmark covering both short-form and long-form scenarios, comprising approximately 28K high-quality evaluation rubrics. We systematically evaluate 12 rerankers: even the best method achieves no more than 45% coverage, cross-document coordination dimensions are universally weak, and no single method maintains top performance across both settings. Building on this, we propose Rubric4Setwise, a training-free method that converts rubric-based evaluation criteria into document set selection signals, achieving the best downstream generation performance with fewer documents and search rounds. It is the only method that maintains state-of-the-art results across both scenarios, validating the effectiveness of closing the loop from evaluation to optimization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。