评估多语言平行数据质量,发现无通用指标,需按语言对定制方案。
Model-Based Quality Assessment for Massively Multilingual Parallel Data

- 用多语言嵌入检测句子是否平行,9种方法测试
- 41,412个语言对中,不同模型表现差异大
- 适合构建多语言数据清洗流水线的研究者
大规模多语言双语语料常存在非平行句对和低质量翻译问题。本文将模型化评估分解为两个独立模块:基于多语言嵌入的平行性判断与无参考质量估计(QE)。在平行性方面,在FLORES-200和BOUQuET检索任务上评测了四种嵌入模型,覆盖6,654个源-目标语言对。在QE方面,对41,412个有序语言对的官方翻译进行九种无参考评估器的测试。结果表明,无模型在所有语言对上均表现可靠;简单集成会稀释强信号,而标注语言覆盖度越高,QE得分也越高。整体说明,多语言平行数据评估应作为面向具体语言对的路由与校准问题处理,不期待单一通用指标适用于所有语言。
原文摘要 · Abstract (English)
Large-scale multilingual bitext often contains two distinct problems: non-parallel sentence pairs and low-quality translations. We decompose model-based assessment for such data into two independent components: parallelism assessment with multilingual embeddings and reference-free quality estimation (QE). For parallelism, we benchmark four embedding models on FLORES-200 and BOUQuET retrieval tasks, covering 6,654 source--target directions in our target language-pair inventory. For QE, we evaluate nine reference-free evaluators on professional FLORES-200 translations across 41,412 ordered source--target directions. Results show that no model is universally reliable across translation directions. Naive QE ensembles dilute strong model signals, while documented target-language coverage is strongly associated with higher QE scores. Overall, these findings suggest that multilingual parallel-data assessment is best approached as a direction-aware routing and calibration problem, where no single universal metric is expected to suffice across all languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。