arXiv:2605.26575cs.CL2026-05

发现跨语言检索不对称主因是‘枢纽效应’,非方向性偏差。

Hubness, Not Anisotropy, Drives Cross-Lingual Retrieval Asymmetry in Multilingual Embedding Models

论文配图:Hubness, Not Anisotropy, Drives Cross-Lingual Retrieval Asymmetry in Multilingual Embedding Models
图 1 · 摘自论文原文
  • 通过五种模型验证,枢纽效应是导致检索不对称的核心机制。
  • 使用改进度量法可解决63.5%的检索不对称问题,效果远超删减向量。
  • 建议用CSLS替代余弦相似度,提升多语言检索公平性。

多语言嵌入模型常假设跨语言检索对称:若语言A的查询能检索到语言B的译文,反之亦然。实际中却常失败。基于包含6,518条英语、孟加拉语、印地语和阿拉伯语习语谚语的平行语料库,使用五种生产级编码器(Gemini、Mistral、OpenAI-L、OpenAI-S、Qwen),我们将其失败形式化为互近邻互异性缺失。在五项预注册实验中,验证单一机制假设:在多语言空间的几何病态中,枢纽效应(hubness)而非方向性(anisotropy)、质心漂移或幅值,是主导因果驱动因素。枢纽质量在联合回归中贡献率达49.5%(1.68倍于次优预测因子;偏决定系数0.302对比方向性的0.003)。引入枢纽感知评分校正(CSLS)可弥补最差与最佳模型间63.5%的互异性差距,其模型内效应量达手术式删除枢纽向量的130倍。该对比揭示机制本质:枢纽效应是相似度度量的病理,而非个别枢纽向量所致。我们通过证明二者统计可分离,解决著名的方向性-枢纽效应悖论,并建议将CSLS作为多语言嵌入流水线的默认检索度量。

原文摘要 · Abstract (English)

Multilingual embedding models are deployed under the assumption that cross-lingual retrieval is symmetric: if a query in language A retrieves its translation in language B, the reverse should also hold. In practice it does not. Using a parallel corpus of 6,518 idiomatic and proverbial expressions in English, Bangla, Hindi, and Arabic, embedded by five production-grade encoders (Gemini, Mistral, OpenAI-L, OpenAI-S, Qwen), we formalise this failure as a deficit in mutual nearest-neighbour reciprocity and test a single mechanistic claim: among the geometric pathologies of multilingual spaces, hubness, not anisotropy, centroid drift, or magnitude, is the dominant causal driver. Across five pre-registered experiments with falsification conditions specified in advance, hub mass dominates a joint regression on reciprocity (49.5% dominance share, 1.68x the next predictor; partial R^2 = 0.302 versus 0.003 for anisotropy), while a hub-aware score correction (CSLS) closes 63.5% of the worst-to-best reciprocity gap and yields a mean within-model effect size 130x larger than surgical hub-vector ablation. The latter contrast pinpoints the mechanism: hubness is a pathology of the similarity metric, not of individual hub vectors. We resolve the well-known anisotropy-hubness paradox by showing the two are statistically dissociable, and we recommend replacing cosine similarity with CSLS as the default retrieval metric for multilingual embedding pipelines.

多语言嵌入检索对称性枢纽效应相似度度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。