arXiv:2509.05425cs.CLcs.AI2025-09

仅凭语言特征就能精准预测翻译质量,无需实际翻译。

No Text Needed: Forecasting MT Quality and Inequity from Fertility and Metadata

  • 用词频比、字数和语言元数据预测翻译质量
  • 对203种语言的GPT-4o翻译预测准确率达R²=0.72
  • 适合做多语言评估与质量估计的研究者参考

我们发现,无需运行翻译系统,仅通过少量特征——词频比、词数及基本语言元数据(语系、书写系统、地区)——即可对FLORES-200基准中203种语言的GPT-4o翻译结果进行高精度的翻译质量预测。梯度提升模型在英文→其他语言的翻译预测中达到R²=0.72,其他语言→英文为R²=0.66。特征重要性分析显示,译入英语时语言类型学因素占主导,而译向多样目标语言时词频比影响更大。研究揭示翻译质量受词级频率与整体语言类型共同塑造,为多语言评估与质量估计提供了新视角。

原文摘要 · Abstract (English)

We show that translation quality can be predicted with surprising accuracy \textit{without ever running the translation system itself}. Using only a handful of features, token fertility ratios, token counts, and basic linguistic metadata (language family, script, and region), we can forecast ChrF scores for GPT-4o translations across 203 languages in the FLORES-200 benchmark. Gradient boosting models achieve favorable performance ($R^{2}=0.66$ for XX$\rightarrow$English and $R^{2}=0.72$ for English$\rightarrow$XX). Feature importance analyses reveal that typological factors dominate predictions into English, while fertility plays a larger role for translations into diverse target languages. These findings suggest that translation quality is shaped by both token-level fertility and broader linguistic typology, offering new insights for multilingual evaluation and quality estimation.

翻译质量无文本预测多语言评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。