arXiv:2604.20763cs.IRcs.AI2026-04被引 1

用语义分层提升检索评估的可信度,解决传统方法的隐藏偏差问题。

Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation

论文配图:Coverage, Not Averages: Semantic Stratification for Trustworthy Retrieval Evaluation
图 1 · 摘自论文原文
  • 基于实体聚类构建可解释的语义空间,系统生成覆盖缺失区域的查询
  • 实现不同检索场景下的语义覆盖保障,发现系统性覆盖率缺口
  • 适合需要透明、稳定评估结果的研究者和工程团队

检索质量是影响检索增强生成(RAG)准确性与鲁棒性的主要瓶颈。当前评估依赖启发式构造的查询集,引入了隐含的内在偏差。本文将检索评估形式化为统计估计问题,指出指标可靠性根本受限于评估集的构建方式。提出“语义分层”方法,通过基于实体的文档聚类,构建可解释的全局语义空间,并系统生成针对缺失分层的查询。该方法实现了(1)跨检索范式的正式语义覆盖保证,(2)对检索失败模式的可解释洞察。在多个基准和检索方法上的实验验证了框架有效性:结果揭示系统性覆盖差距,识别出解释检索性能方差的结构性信号,表明分层评估比平均指标更稳定、透明,支持更可信的决策。

原文摘要 · Abstract (English)

Retrieval quality is the primary bottleneck for accuracy and robustness in retrieval-augmented generation (RAG). Current evaluation relies on heuristically constructed query sets, which introduce a hidden intrinsic bias. We formalize retrieval evaluation as a statistical estimation problem, showing that metric reliability is fundamentally limited by the evaluation-set construction. We further introduce \emph{semantic stratification}, which grounds evaluation in corpus structure by organizing documents into an interpretable global space of entity-based clusters and systematically generating queries for missing strata. This yields (1) formal semantic coverage guarantees across retrieval regimes and (2) interpretable visibility into retrieval failure modes. Experiments across multiple benchmarks and retrieval methods validate our framework. The results expose systematic coverage gaps, identify structural signals that explain variance in retrieval performance, and show that stratified evaluation yields more stable and transparent assessments while supporting more trustworthy decision-making than aggregate metrics.

检索评估语义分层RAG可信性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。