arXiv:2511.05684cs.IRcs.CL2025-11Conference of the …被引 1

无需训练,提升零样本文档检索的区分能力

A Representation Sharpening Framework for Zero Shot Dense Retrieval

  • 通过增强文档表示,区分语义相近的文档
  • 在二十多个多语言数据集上超越传统方法,刷新BRIGHT基准
  • 兼容已有方法,索引时近似可零推理开销保持性能

零样本密集检索面临无相关查询的问题,依赖预训练的密集检索器(DRs),但这些模型未在目标语料库上训练,难以区分语义相似文档。为此,我们提出一种无需训练的表示锐化框架,在不改变原有模型的前提下,通过补充信息增强文档表示,帮助区分相似文档。在超过二十个跨语言数据集上,该框架始终优于传统检索,创下BRIGHT基准新纪录。我们证明其与已有零样本密集检索方法兼容,能持续提升性能。此外,针对性能与成本权衡问题,我们设计了索引时近似方案,在几乎不增加推理开销的情况下保留主要性能优势。

原文摘要 · Abstract (English)

Zero-shot dense retrieval is a challenging setting where a document corpus is provided without relevant queries, necessitating a reliance on pretrained dense retrievers (DRs). However, since these DRs are not trained on the target corpus, they struggle to represent semantic differences between similar documents. To address this failing, we introduce a training-free representation sharpening framework that augments a document's representation with information that helps differentiate it from similar documents in the corpus. On over twenty datasets spanning multiple languages, the representation sharpening framework proves consistently superior to traditional retrieval, setting a new state-of-the-art on the BRIGHT benchmark. We show that representation sharpening is compatible with prior approaches to zero-shot dense retrieval and consistently improves their performance. Finally, we address the performance-cost tradeoff presented by our framework and devise an indexing-time approximation that preserves the majority of our performance gains over traditional retrieval, yet suffers no additional inference-time cost.

零样本检索表示增强密集检索多语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。