arXiv:2501.14434cs.IRcs.LG2025-01被引 1

通过重采样难负例提升生成伪标签的域适应效果

Remining Hard Negatives for Generative Pseudo Labeled Domain Adaptation

  • 在知识蒸馏中动态重采样难负例以优化检索性能
  • 在BEIR和LoTTe数据集上分别提升13/14和9/12项指标
  • 适合需要跨域检索鲁棒性的信息检索研究者

密集检索器在神经信息检索中展现出巨大潜力,但在面对领域偏移时缺乏鲁棒性,限制了其在零样本跨领域场景下的表现。当前最先进的域适应技术是生成伪标签(GPL),它利用合成查询生成和初始挖掘的难负例,将交叉编码器的知识蒸馏到目标领域的密集检索器中。本文分析了域适应模型返回的文档,发现其相较于非域适应模型更贴近目标查询。为此,我们提出在知识蒸馏阶段重新构建难负例索引,以挖掘更优的难负例。所提出的R-GPL方法在14个BEIR数据集中的13个和12个LoTTe数据集中的9个上提升了排序性能。主要贡献包括:(i) 分析域适应与非域适应模型返回的难负例差异;(ii) 在LoTTe和BEIR数据集上对比了含与不含难负例重采样的GPL训练效果。

原文摘要 · Abstract (English)

Dense retrievers have demonstrated significant potential for neural information retrieval; however, they exhibit a lack of robustness to domain shifts, thereby limiting their efficacy in zero-shot settings across diverse domains. A state-of-the-art domain adaptation technique is Generative Pseudo Labeling (GPL). GPL uses synthetic query generation and initially mined hard negatives to distill knowledge from cross-encoder to dense retrievers in the target domain. In this paper, we analyze the documents retrieved by the domain-adapted model and discover that these are more relevant to the target queries than those of the non-domain-adapted model. We then propose refreshing the hard-negative index during the knowledge distillation phase to mine better hard negatives. Our remining R-GPL approach boosts ranking performance in 13/14 BEIR datasets and 9/12 LoTTe datasets. Our contributions are (i) analyzing hard negatives returned by domain-adapted and non-domain-adapted models and (ii) applying the GPL training with and without hard-negative re-mining in LoTTE and BEIR datasets.

域适应检索增强知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。