arXiv:2604.01957cs.CLcs.IR2026-04中稿 · LREC 2026

自动化检测20种语言翻译数据集质量,发现错误分布不均。

Diagnosing Translated Benchmarks: An Automated Quality Assurance Study of the EU20 Benchmark Suite

  • 用三步法自动评估欧盟20语种基准集翻译质量
  • COMET分数越低,片段级错误越多(如HellaSwag)
  • 适合需高效验证多语言数据的NLP研究者

机器翻译基准数据集可降低成本并实现规模扩展,但噪声、结构丢失和质量不均会削弱可信度。关键问题不仅在于能否翻译,更在于能否规模化测量与验证翻译可靠性。本研究针对包含五个基准测试、翻译为20种语言的EU20基准套件,采用三步自动化质量保障方法:(i) 结构化语料库审计并实施针对性修正;(ii) 利用神经评估指标COMET(无参考与有参考)对比DeepL、ChatGPT、Google翻译服务的质量;(iii) 基于大模型构建片段级翻译错误图谱。结果一致显示:COMET分数较低的数据集在片段层面存在更高比例的准确率/误译错误(尤其HellaSwag);MMLU上基于参考的COMET结果与人工编辑样本对比也支持该趋势。研究发布清洗/修正后的EU20数据集及可复现代码。总体而言,自动化质量保障提供实用、可扩展的指标,有助于优先排序人工审核——补充而非替代人工黄金标准。

原文摘要 · Abstract (English)

Machine-translated benchmark datasets reduce costs and offer scale, but noise, loss of structure, and uneven quality weaken confidence. What matters is not merely whether we can translate, but also whether we can measure and verify translation reliability at scale. We study translation quality in the EU20 benchmark suite, which comprises five established benchmarks translated into 20 languages, via a three-step automated quality assurance approach: (i) a structural corpus audit with targeted fixes; (ii) quality profiling using a neural metric (COMET, reference-free and reference-based) with translation service comparisons (DeepL / ChatGPT / Google); and (iii) an LLM-based span-level translation error landscape. Trends are consistent: datasets with lower COMET scores exhibit a higher share of accuracy/mistranslation errors at span level (notably HellaSwag; ARC is comparatively clean). Reference-based COMET on MMLU against human-edited samples points in the same direction. We release cleaned/corrected versions of the EU20 datasets, and code for reproducibility. In sum, automated quality assurance offers practical, scalable indicators that help prioritize review -- complementing, not replacing, human gold standards.

多语言质量评估自动化基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。