arXiv:2602.20800cs.IR2026-02

通过分离评价与监督模型,避免偏好泄露,提升文化相关性排序的可信度。

Mitigating Preference Leakage via Strict Estimator Separation for Normative Generative Ranking

  • 构建双评委框架,严格分离训练与评估模型,防止偏好泄露。
  • 在3.3万条文化故事数据上,蒸馏模型性能远超原始交叉编码器。
  • 适用于需要公平评估生成式检索中文化偏好的研究者。

在生成式信息检索(GenIR)中,候选选择已成为瓶颈,尤其针对文化相关性等规范性标准。当前基于大模型作为裁判的评估常因循环性和偏好泄露导致性能虚高,即训练与评估模型重叠。本文将文化相关性形式化为查询内排序任务,提出无泄露的双评委框架,严格分离监督(裁判B)与评估(裁判A)模型。在包含33,052个文化背景故事的新基准(NGR-33k)上,经典基线仅带来有限提升;而由裁判B监督的交叉编码器蒸馏得到的密集双编码器(BGE-M3),在无泄露裁判A评估下表现显著更优。该框架在人工标注的道德故事数据集上也展现出与人类规范的高度一致性。结果表明,严格的评估者分离是可信GenIR评估的前提,证明细微文化偏好可被有效提炼为高效排序器而无泄露风险。

原文摘要 · Abstract (English)

In Generative Information Retrieval (GenIR), the bottleneck has shifted from generation to the selection of candidates, particularly for normative criteria such as cultural relevance. Current LLM-as-a-Judge evaluations often suffer from circularity and preference leakage, where overlapping supervision and evaluation models inflate performance. We address this by formalising cultural relevance as a within-query ranking task and introducing a leakage-free two-judge framework that strictly separates supervision (Judge B) from evaluation (Judge A). On a new benchmark of 33,052 (NGR-33k) culturally grounded stories, we find that while classical baselines yield only modest gains, a dense bi-encoder distilled from a Judge-B-supervised Cross-Encoder is highly effective. Although the Cross-Encoder provides a strong supervision signal for distillation, the distilled BGE-M3 model substantially outperforms it under leakage-free Judge~A evaluation. We validate our framework on the human-curated Moral Stories dataset, showing strong alignment with human norms. Our results demonstrate that rigorous evaluator separation is a prerequisite for credible GenIR evaluation, proving that subtle cultural preferences can be distilled into efficient rankers without leakage.

生成式检索文化相关性模型蒸馏评估可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。