探究密集检索器领域自适应性能的关键影响因素
On Correlating Factors for Domain Adaptation Performance
- 分析生成查询分布对领域自适应的影响
- 生成与目标域相似的查询可显著提升性能
- 强调应针对具体领域定制生成查询
密集检索器在神经信息检索中展现出巨大潜力,但在面对领域迁移时缺乏鲁棒性,限制了其在跨领域零样本场景下的有效性。本文旨在分析导致密集检索器领域自适应成功的关键因素。我们引入源域与生成查询之间的领域相似性代理指标,并通过案例研究对比两种强大的领域自适应技术。结果发现,生成查询的类型分布是关键因素;若生成的查询与测试文档所属领域相近,则能有效提升领域自适应方法的表现。本研究进一步强调了为特定领域定制生成查询的重要性。
原文摘要 · Abstract (English)
Dense retrievers have demonstrated significant potential for neural information retrieval; however, they lack robustness to domain shifts, limiting their efficacy in zero-shot settings across diverse domains. In this paper, we set out to analyze the possible factors that lead to successful domain adaptation of dense retrievers. We include domain similarity proxies between generated queries to test and source domains. Furthermore, we conduct a case study comparing two powerful domain adaptation techniques. We find that generated query type distribution is an important factor, and generating queries that share a similar domain to the test documents improves the performance of domain adaptation methods. This study further emphasizes the importance of domain-tailored generated queries.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。