arXiv:2608.28776cs.CL2026-08

对比五种多语言嵌入模型,发现它们能识别翻译中的语义错误,但需结合其他方法使用。

Evaluating Multilingual Sentence Embeddings for Translation Error Detection:An English--Greek Contrastive Study

  • 构建英文-希腊语对照数据集,覆盖15类翻译错误,用于评估嵌入模型敏感性。
  • BGE-M3准确率达89.3%,优于其他嵌入模型,但低于基准模型COMETKiwi的94.49%。
  • 嵌入模型对事实和词汇错误更敏感,而对时态、代词等错误检测效果较差。

多语言句子嵌入被广泛用于跨语言语义相似性估计,但其对细微翻译错误的敏感性仍不明确。本研究考察通用多语言嵌入模型能否区分正确与轻微修改的英-希翻译。基于FLORES+句对参考译文,由两名译者审核构建了包含1,850个样本的对比数据集,涵盖10类核心与5类探索性错误,涉及事实、词义、语法、关系、指代及话语层面现象。评估了五种多语言句子嵌入模型(BGE-M3、Multilingual E5、Multilingual MPNet、LaBSE、Jina Embeddings v3)在源句与正确/错误译文间余弦相似度的表现,并以无参考的COMETKiwi作为机器翻译质量评估基线。通过对比准确率和得分差值评估类别敏感性。结果显示,BGE-M3在嵌入模型中表现最佳,准确率为89.30%,而COMETKiwi达94.49%。嵌入模型对显性事实与词汇变化检测更可靠,但对时态-体态及代词-指代错误敏感度较低;COMETKiwi在时态、代词、语义角色等复杂类别上表现更优,但在日期时间与数字错误上敏感度低,且在数字相关错误上劣于嵌入模型。结果表明,多语言嵌入提供有用语义充分性信号,但更适合作为综合翻译评估框架的组成部分,而非独立指标。

原文摘要 · Abstract (English)

Multilingual sentence embeddings are increasingly used to estimate semantic similarity across languages, yet their sensitivity to fine-grained translation errors remains insufficiently understood. This study investigates whether general-purpose multilingual embedding models can distinguish correct English-Greek translations from minimally modified erroneous alternatives. A contrastive dataset was developed from FLORES+ sentence-aligned reference translations and reviewed by two translation experts. It contains 1,850 examples across ten core and five exploratory error categories, covering factual, lexical-semantic, grammatical, relational, referential, and discourse-level phenomena. Five multilingual sentence-embedding models (BGE-M3, Multilingual E5, Multilingual MPNet, LaBSE, and Jina Embeddings v3) were evaluated using cosine similarity between each English source sentence and its correct and erroneous Greek translations. A reference-free COMETKiwi model was also evaluated as an MT quality-estimation baseline. Performance was assessed through contrastive accuracy and score margins for category-specific sensitivity. BGE-M3 achieved the highest accuracy among embedding models at 89.30 percent, while COMETKiwi achieved 94.49 percent. Embedding models detected explicit factual and lexical changes more reliably than tense-and-aspect and pronoun-coreference errors. COMETKiwi improved performance on several difficult categories, including tense and aspect, pronoun and coreference, and semantic-role errors, but showed lower sensitivity to date-and-time errors and underperformed the embedding models on numbers. The results show complementary error-sensitivity profiles: multilingual sentence embeddings provide useful semantic adequacy signals but are better suited as components of broader translation-evaluation frameworks than as standalone metrics.

多语言嵌入翻译评估误差检测语义相似性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。