发现翻译错误是多语言大模型评估不准的主因
Quantifying the Impact of Translation Errors on Multilingual LLM Evaluation

- 用LLM和自动质检模型检测翻译错位,比对人工标注
- 翻译错误导致准确率明显下降,即使英文原题正确
- 适合关注多语言评估可靠性的研究者
机器翻译的评测集被广泛用于评估大语言模型的多语言能力,但其中的翻译错误尚未得到充分研究,影响了多语言评估的可靠性与可比性。本文填补两个实际空白:(i) 比较大语言模型判别器与基于跨度感知的自动翻译质量评估模型(xCOMET-XXL)在自然生成的评测集翻译中,与人工专家标注的错位段落的一致性;(ii) 探究翻译错误(而非源语言侧问题)对翻译后评测准确率下降的影响程度。结果表明,自然翻译中段落级错位标注的一致性并不理想,且目标语言侧的翻译错误始终与可观测的准确率下降相关,即使在控制英语原题正确性和源端异常后依然显著。
原文摘要 · Abstract (English)
Machine-translated benchmarks are widely used to assess the multilingual capabilities of large language models (LLMs), yet translation errors in these benchmarks remain underexplored, raising concerns about the reliability and comparability of multilingual evaluation. We address two practical gaps: (i) how well automatic MQM-style error spans from LLM judges and a span-aware QE baseline (xCOMET-XXL) match expert human span annotations on benchmark translations, and (ii) how strongly translation errors (as opposed to source-side issues in the English original) explain accuracy drops on translated benchmarks. We find that span agreement is non-trivial on naturally occurring benchmark translations, and that target-side translation errors are consistently associated with measurable, percentage-point drops in translated accuracy even after controlling for English correctness and source-side anomalies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。