arXiv:2603.18005cs.IR2026-03中稿 · findings EACL 2026综述

综述密集检索中负样本选择技术,涵盖从随机到大模型生成的新方法。

Negative Sampling Techniques in Information Retrieval: A Survey

  • 按来源将负样本技术分为随机、静态/动态挖掘和合成数据三类
  • 对比分析不同方法在效果、计算开销和实现难度间的权衡
  • 重点覆盖大模型生成合成数据的前沿方向,适合检索系统研究者

信息检索(IR)是众多现代自然语言处理应用的基础。密集检索(DR)利用神经网络学习语义向量表示,显著提升了检索性能。训练有效的密集检索器依赖对比学习中的高质量负样本选择。本文综合35篇奠基性论文,全面梳理密集信息检索中的负样本技术。独特贡献在于聚焦现代NLP应用,并纳入近期大语言模型(LLM)驱动的生成方法,填补了以往综述的空白。我们提出一个分类体系,将技术分为随机、静态/动态挖掘及合成数据三类,并从有效性、计算成本与实现难度三方面进行分析。最后指出当前挑战与未来方向,特别是大模型生成合成数据的应用潜力。

原文摘要 · Abstract (English)

Information Retrieval (IR) is fundamental to many modern NLP applications. The rise of dense retrieval (DR), using neural networks to learn semantic vector representations, has significantly advanced IR performance. Central to training effective dense retrievers through contrastive learning is the selection of informative negative samples. Synthesizing 35 seminal papers, this survey provides a comprehensive and up-to-date overview of negative sampling techniques in dense IR. Our unique contribution is the focus on modern NLP applications and the inclusion of recent Large Language Model (LLM)-driven methods, an area absent in prior reviews. We propose a taxonomy that categorizes techniques including random, static/dynamically mined, and synthetic datasets. We then analyze these approaches with respect to trade-offs between effectiveness, computational cost, and implementation difficulty. The survey concludes by outlining current challenges and promising future directions for the use of LLM-generated synthetic data.

信息检索负样本密集检索大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。