arXiv:2506.21581cs.IRcs.AI2025-06

不同评估基准让同一模型的改进效果看起来差3.6倍,选对测试方法很重要。

Evaluating the Robustness of Dense Retrievers in Interdisciplinary Domains

  • 用两个语义结构不同的基准测试,发现微调效果差异巨大。
  • 相同方法在重叠语义场景下提升达2.22% NDCG,分离场景仅0.61%。
  • 适合关注跨领域检索系统评估的开发者和研究者。

评估基准的特性会扭曲领域适配在检索模型中的真实收益,导致部署决策误导。我们以环境监管文件检索为例,对ColBERTv2模型在联邦机构环境影响声明(EIS)上进行微调,并在两个具有不同语义结构的基准上评估。结果表明,相同领域适配方法在不同评估框架下呈现截然不同的表现:在一个主题边界清晰的基准上,性能提升最大仅0.61% NDCG;而在一个语义重叠的基准上,提升高达2.22% NDCG,相差3.6倍。通过主题多样性指标分析,高表现基准的上下文间平均余弦距离高11%,轮廓系数低23%,直接解释了性能差异。这说明评估框架的选择强烈影响对检索系统有效性的判断。主题分离的基准常低估领域适配收益,而具有重叠语义边界的基准更能反映真实监管文档的复杂性。该发现对跨领域AI系统的开发与部署具有重要启示。

原文摘要 · Abstract (English)

Evaluation benchmark characteristics may distort the true benefits of domain adaptation in retrieval models. This creates misleading assessments that influence deployment decisions in specialized domains. We show that two benchmarks with drastically different features such as topic diversity, boundary overlap, and semantic complexity can influence the perceived benefits of fine-tuning. Using environmental regulatory document retrieval as a case study, we fine-tune ColBERTv2 model on Environmental Impact Statements (EIS) from federal agencies. We evaluate these models across two benchmarks with different semantic structures. Our findings reveal that identical domain adaptation approaches show very different perceived benefits depending on evaluation methodology. On one benchmark, with clearly separated topic boundaries, domain adaptation shows small improvements (maximum 0.61% NDCG gain). However, on the other benchmark with overlapping semantic structures, the same models demonstrate large improvements (up to 2.22% NDCG gain), a 3.6-fold difference in the performance benefit. We compare these benchmarks through topic diversity metrics, finding that the higher-performing benchmark shows 11% higher average cosine distances between contexts and 23% lower silhouette scores, directly contributing to the observed performance difference. These results demonstrate that benchmark selection strongly determines assessments of retrieval system effectiveness in specialized domains. Evaluation frameworks with well-separated topics regularly underestimate domain adaptation benefits, while those with overlapping semantic boundaries reveal improvements that better reflect real-world regulatory document complexity. Our findings have important implications for developing and deploying AI systems for interdisciplinary domains that integrate multiple topics.

检索评估领域适配跨领域

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。