arXiv:2602.21543cs.CLcs.AI2026-02

用多语言平行文本提升跨语言嵌入对齐效果

Enhancing Multilingual Embeddings via Multi-Way Parallel Text Alignment

  • 基于多语言平行语料,通过对比学习增强跨语言对齐
  • 在多个任务上实现21.3%~28.4%的性能提升,覆盖已见与未见语言
  • 适用于需要高质量跨语言表示的NLU任务,尤其适合小数据场景

多语言预训练通常缺乏显式对齐信号,导致表示空间中的跨语言对齐效果不佳。本文表明,使用包含六种目标语言的多语言平行语料库,对标准预训练模型进行跨语言对齐训练,可显著提升多语言和跨语言表示能力。我们利用现成的神经机器翻译模型生成英文文本在六种目标语言上的翻译,构建多向平行数据集,并通过对比学习实现强跨语言对齐。该方法在MTEB基准上对XLM-Roberta和多语言BERT base模型均带来显著提升,相比以英语为中心的双语平行数据(En-X),在双语挖掘(+21.3%)、语义相似度(+5.3%)和分类任务(+28.4%)上表现更优。此外,在小规模数据上微调mE5模型时,引入多向平行性显著提升双语挖掘效果,证明即使已有高质量句向量预训练模型,多向跨语言监督仍至关重要。

原文摘要 · Abstract (English)

Multilingual pretraining typically lacks explicit alignment signals, leading to suboptimal cross-lingual alignment in the representation space. In this work, we show that training standard pretrained models for cross-lingual alignment with a multi-way parallel corpus in a diverse pool of languages can substantially improve multilingual and cross-lingual representations for NLU tasks. We construct a multi-way parallel dataset using translations of English text from an off-the-shelf NMT model for a pool of six target languages and achieve strong cross-lingual alignment through contrastive learning. This leads to substantial performance gains across both seen and unseen languages for multiple tasks from the MTEB benchmark evaluated for XLM-Roberta and multilingual BERT base models. Using a multi-way parallel corpus for contrastive training yields substantial gains on bitext mining (21.3%), semantic similarity (5.3%), and classification (28.4%) compared to English-centric (En-X) bilingually parallel data, where X is sampled from a pool of multiple target languages. Furthermore, finetuning mE5 model on a small dataset with multi-way parallelism significantly improves bitext mining compared to one without, underscoring the importance of multi-way cross-lingual supervision even for models already pretrained for high-quality sentence embeddings.

多语言嵌入对比学习跨语言对齐NLU

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。