arXiv:2607.00171cs.CL2026-07

用英语最小差异对评估任意语言的嵌入表示,解决多语言评估不均衡问题。

ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs

论文配图:ALEE: Any-Language Evaluation of Embeddings via English-Centric Minimal Pairs
图 1 · 摘自论文原文
  • 基于语义最小差异对生成跨语言测试样本,控制语义变化粒度。
  • 在275+语言上测试发现模型表现差异显著,与训练数据覆盖度相关。
  • 适合研究多语言模型、低资源语言表征或评估嵌入鲁棒性者使用。

文本嵌入广泛用于语义相似性任务,但其评估仍面临挑战:现有基准静态、覆盖语言有限、常受领域偏差影响,且难以代表低资源语言。为此,我们提出ALEE,将Sentence Smith(Li et al., 2025)扩展至跨语言与段落级别。ALEE利用抽象意义表示(AMR)生成具有精细语义变化的英语最小差异对,并配以目标语言的平行翻译,实现对任意语言嵌入模型的精准诊断。我们在三个并行数据集上对275+种语言和多种嵌入模型进行了大规模实证研究。结果显示,不同语言、文本长度及语言现象下性能差异显著,暴露了跨语言语义表示的持续缺陷,这些缺陷与训练数据中语言的流行程度及子词分词方式密切相关。代码已开源:https://github.com/Andrian0s/any-lang-embed-eval

原文摘要 · Abstract (English)

Text embeddings are standard for semantic similarity tasks, yet their evaluation remains an open challenge. Current benchmarks are static, cover only a limited set of languages, are often domain-specific, susceptible to overfitting, and poorly representative of low-resource languages. To address these limitations, we introduce ALEE, a framework that extends Sentence Smith (Li et al., 2025) to the cross-lingual and paragraph level. ALEE uses Abstract Meaning Representations (AMR) to generate English minimal pairs with controlled, fine-grained semantic shifts, which are paired with translations in target languages. This approach enables targeted diagnostics for models in any language with English parallel data. We conduct a large-scale empirical study across a diverse set of embedding models and 275+ languages spanning three parallel datasets. On ALEE, performance varies substantially across languages, text lengths, and linguistic phenomena, exposing persistent gaps in cross-lingual semantic representation that track language prevalence in training resources and subword tokenization. We release ALEE at https://github.com/Andrian0s/any-lang-embed-eval

嵌入评估多语言最小差异对语义表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。