检验机器翻译是否保持文本语义相似性,发现10种语言不变、4种明显失真。
Is Textual Similarity Invariant under Machine Translation? Evidence Based on the Political Manifesto Corpus

- 用嵌入模型间一致性判断翻译对语义关系的影响
- 在28语言2800+政纲数据上识别出10种保持不变的语言
- 方法可推广至其他任务,适合多语言语义研究者
我们研究了段落嵌入间的余弦相似性在机器翻译下的不变性,基于包含2800多个政治党派纲领的《政纲语料库》,这些文本通过欧盟eTranslation服务被译成英文。不直接测量翻译引起的语义偏移,而是通过嵌入模型间对原始语言文本的一致性差异,设定不变性阈值。该框架对四种关于翻译与嵌入选择交互作用的假设进行逐语言非劣性检验,能区分出翻译明显保持语义结构的语言、明显破坏结构的语言,以及证据不足无法判断的语言。该方法与语料和处理流程无关,可自然拓展至下游任务。应用于本数据,识别出10种具有翻译不变性的语言,4种存在可检测到的语义扭曲。
原文摘要 · Abstract (English)
We investigate the extent to which cosine similarity between paragraph embeddings is invariant under machine translation, using the Manifesto Corpus of over 2,800 political party platforms in 28 languages translated to English via the EU eTranslation service. Rather than measuring translation-induced semantic shift directly we measure the stability of pairwise similarity relationships across embedding models, and use inter-model disagreement on original-language text as a calibrated invariance threshold. This yields a per-language non-inferiority test for four hypotheses about how translation interacts with embedding choice, with verdicts that distinguish languages where translation demonstrably preserves semantic structure from those where it demonstrably degrades it and from those where the available evidence does not resolve the question. The framework is corpus- and pipeline-agnostic and extends naturally to downstream tasks. Applied to our data, it identifies ten languages with translation invariance and four with detectable distortion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。