首个系统诊断文档位置偏见的基准,覆盖多语言多领域。
PosIR: Position-Aware Heterogeneous Information Retrieval Benchmark
- 按文档长度分桶,分离位置与长度影响,精准诊断偏见。
- 长文档(>1536词)上模型表现与现有评测相关性低。
- 多数模型有首段偏好,部分出现末段偏好,机制可追溯。
真实文档中信息可能分布在任意位置,导致检索模型存在位置偏见——即因内容位置不同而系统性偏好或忽略。尽管已有研究指出此问题,但现有分析主要集中于英文、未区分文档长度与位置因素,且缺乏标准化诊断框架。为此,我们提出PosIR(Position-Aware Information Retrieval),首个系统化诊断位置偏见的基准。它包含310个数据集,覆盖10种语言和31个领域,相关性基于精确参考片段标注。核心方法采用长度控制的分桶策略:按正样本文档长度分组,在每组内分析位置效应,严格隔离位置偏见与长度带来的性能下降。对10个先进嵌入式检索模型的实验表明:(1) 文档超过1536词时,模型在PosIR上的表现与MMTEB基准相关性差,暴露当前短文本评测的局限;(2) 位置偏见普遍存在,且随文档变长加剧,多数模型呈现首段偏好(primacy bias),部分模型显示意外末段偏好(recency bias);(3) 探索性梯度显著性分析揭示两种与位置偏好相关的内部机制。我们希望PosIR能成为推动位置鲁棒检索系统发展的关键诊断工具。
原文摘要 · Abstract (English)
In real-world documents, the information relevant to a user query may reside anywhere from the beginning to the end. This makes position bias -- a systematic tendency of retrieval models to favor or neglect content based on its location -- a critical concern. Although recent studies have identified such bias, existing analyses focus predominantly on English, fail to disentangle document length from information position, and lack a standardized framework for systematic diagnosis. To address these limitations, we introduce PosIR (Position-Aware Information Retrieval), the first standardized benchmark designed to systematically diagnose position bias in diverse retrieval scenarios. PosIR comprises 310 datasets spanning 10 languages and 31 domains, with relevance tied to precise reference spans. At its methodological core, PosIR employs a length-controlled bucketing strategy that groups queries by positive document length and analyzes positional effects within each bucket. This design strictly isolates position bias from length-induced performance degradation. Extensive experiments on 10 state-of-the-art embedding-based retrieval models reveal that: (1) retrieval performance on PosIR with documents exceeding 1536 tokens correlates poorly with the MMTEB benchmark, exposing limitations of current short-text evaluations; (2) position bias is pervasive in embedding models and even increases with document length, with most models exhibiting primacy bias while certain models show unexpected recency bias; (3) as an exploratory investigation, gradient-based saliency analysis further uncovers two distinct internal mechanisms that correlate with these positional preferences. We hope that PosIR can serve as a valuable diagnostic framework to advance the development of position-robust retrieval systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。