arXiv:2606.13537cs.CL2026-06ACL

研究多语言查询混合效果,发现英语主导下混合查询有规律可循。

When Does Mixing Help? Analyzing Query Embedding Interpolation in Multilingual Dense Retrieval

论文配图:When Does Mixing Help? Analyzing Query Embedding Interpolation in Multilingual Dense Retrieval
图 1 · 摘自论文原文
  • 通过嵌入向量插值构造混合查询,系统测试不同语种比例下的检索效果。
  • 88/105 情况下混合查询优于单一语言查询,非英文档库中混合更有效。
  • 英语是最佳混合伙伴,且语言差异越大,混合收益越低。

尽管跨语言查询在多语言社区中普遍存在,但稠密检索器对这类查询的敏感性仍不明确。本文在 mMARCO 数据集上开展比例控制实验,通过嵌入级混合方式系统评估不同语种查询翻译比例下的检索性能——将混合查询构建为单语嵌入的插值。基于 BGE-M3 的实验表明,在 88/105 的情况下,最优混合比例优于表现最好的单语基准。研究发现存在显著的英语主导不对称性:在非英语文档索引中混合始终有益,而含英语文档的索引则以纯英语查询为佳。此外,英语对所有非英语文档语言均为最强混合伙伴。当控制英语主导效应后,混合增益与语言类型学距离呈负相关。结论表明语言混合敏感性具有结构性和可预测性,且在不同模型族与规模下均具鲁棒性。

原文摘要 · Abstract (English)

While mixed-language querying is ubiquitous in multilingual communities, the sensitivity of dense retrievers to such queries remains poorly understood. We present a ratio-controlled study on mMARCO that systematically evaluates retrieval performance by varying the mixing proportion of parallel query translations via embedding-level mixing -- constructing mixed queries as an interpolation of monolingual embeddings. Experiments with BGE-M3 demonstrate that an optimal mixing ratio outperforms the best monolingual endpoint in 88/105 cases. We uncover a distinct asymmetry driven by English dominance: mixing is uniformly beneficial when retrieving from non-English document indices, whereas indices containing English are best served by pure English queries. Furthermore, English acts as the strongest mixing partner for every non-English document language. Finally, when controlling for English dominance, mixing gains correlate negatively with typological distance. We conclude that language-mix sensitivity is structured and predictable, and we validate the robustness of these patterns across model families and scales.

多语言检索嵌入混合语言不对称BGE-M3

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。