仅凭语言特征就能精准预测翻译质量,无需实际翻译。
No Text Needed: Forecasting MT Quality and Inequity from Fertility and Metadata
- 用词频比、字数和语言元数据预测翻译质量
- 对203种语言的GPT-4o翻译预测准确率达R²=0.72
- 适合做多语言评估与质量估计的研究者参考
我们发现,无需运行翻译系统,仅通过少量特征——词频比、词数及基本语言元数据(语系、书写系统、地区)——即可对FLORES-200基准中203种语言的GPT-4o翻译结果进行高精度的翻译质量预测。梯度提升模型在英文→其他语言的翻译预测中达到R²=0.72,其他语言→英文为R²=0.66。特征重要性分析显示,译入英语时语言类型学因素占主导,而译向多样目标语言时词频比影响更大。研究揭示翻译质量受词级频率与整体语言类型共同塑造,为多语言评估与质量估计提供了新视角。
原文摘要 · Abstract (English)
We show that translation quality can be predicted with surprising accuracy \textit{without ever running the translation system itself}. Using only a handful of features, token fertility ratios, token counts, and basic linguistic metadata (language family, script, and region), we can forecast ChrF scores for GPT-4o translations across 203 languages in the FLORES-200 benchmark. Gradient boosting models achieve favorable performance ($R^{2}=0.66$ for XX$\rightarrow$English and $R^{2}=0.72$ for English$\rightarrow$XX). Feature importance analyses reveal that typological factors dominate predictions into English, while fertility plays a larger role for translations into diverse target languages. These findings suggest that translation quality is shaped by both token-level fertility and broader linguistic typology, offering new insights for multilingual evaluation and quality estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。