arXiv:2608.02189cs.IRcs.CL2026-08

分离语义与语言特征,提升多语言检索的零样本迁移能力

Disentangled Contrastive Learning for Zero-Shot Multilingual Dense Retrieval

论文配图:Disentangled Contrastive Learning for Zero-Shot Multilingual Dense Retrieval
图 1 · 摘自论文原文
  • 通过解耦语义与语言子空间,减少跨语言干扰
  • 在mMARCO和MIRACL上超越多个强基线模型
  • 适合需要低资源语言检索的场景

多语言稠密检索旨在基于统一检索模型处理跨语言查询与文档。其挑战在于如何在标注数据稀缺的低资源语言中实现鲁棒的检索迁移。尽管已有研究将高资源语言监督迁移到低资源语言,但共享表示常纠缠语义与语言特征,影响语义相关性优化。不同于以往在纠缠状态下学习语言无关语义特征的方法,本文提出解耦对比学习(DCL)用于多语言稠密检索,将多语言表示分解为语义与语言子空间。设计基于层次化语义对齐和语言去偏对比学习的解耦优化目标,在句子与词级别对齐跨语言语义,同时在语言子空间捕捉语言特异性差异,降低语言干扰。联合优化检索目标,促进英语监督到多语言检索的稳定零样本迁移。在mMARCO和MIRACL上的大量实验表明,该方法持续优于多个强基线,验证了其有效性与泛化能力。

原文摘要 · Abstract (English)

Multilingual dense retrieval aims to handle queries and documents across different languages based on a unified retriever model. The challenge lies in enabling robust retrieval transfer to low-resource languages where annotated retrieval data is often scarce. Although previous studies transfer high-resource supervision to low-resource languages in multilingual semantic representation learning, the shared representation often entangles semantic and linguistic features, which may interfere with optimizing semantic relevance for retrieval. Different from existing methods that focus on learning language-agnostic semantic features under such entanglement, we propose a disentangled contrastive learning~(DCL) method for multilingual dense retrieval by separating multilingual representations into semantic and linguistic subspaces. Specifically, we design disentangled optimization objectives based on hierarchical semantic alignment and language debiasing contrastive learning. By aligning retrieval-relevant semantics across languages at both sentence and token levels while capturing language-specific variations in the linguistic subspace, these objectives reduce language-induced interference in semantic matching. We jointly optimize them with the retrieval objective to facilitate stable zero-shot transfer from English supervision to multilingual dense retrieval. Extensive experiments on mMARCO and MIRACL show that our method consistently outperforms several strong baselines, demonstrating its effectiveness and generalization ability.

多语言检索对比学习解耦表征零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。