提出评估极低资源翻译的新指标,揭示性能差异主因是数据重叠而非模型能力。
Translation or Recitation? Calibrating Evaluation Scores for Machine Translation of Extremely Low-Resource Languages
- 引入FRED四维指标,量化数据集内在难度
- 发现性能差异主要由训练测试重叠和预训练暴露决定
- 指出濒危语言因分词覆盖差导致模型迁移受限
极低资源机器翻译领域报告的性能波动令人困惑,难以跨语言对比较。为解决此问题,我们提出FRED难度指标体系,包含词频比(F)、检索代理(R)、预训练暴露(E)和语料多样性(D),作为数据集固有指标以解释性能得分。分析显示,结果差异主要源于训练-测试数据重叠与预训练暴露程度,而非模型能力。此外,我们发现包括已灭绝及非拉丁字母原住民语言在内的部分语言存在严重分词覆盖率不足(高词频比),暴露出从高资源语言迁移模型时缺乏共享词汇的根本局限。通过在性能评分外提供这些指数,我们推动了跨语言迁移评估的透明化,并为极低资源翻译研究社区建立更可靠的基准。
原文摘要 · Abstract (English)
The landscape of extremely low-resource machine translation (MT) is characterized by perplexing variability in reported performance, often making results across different language pairs difficult to contextualize. For researchers focused on specific language groups -- such as ancient languages -- it is nearly impossible to determine if breakthroughs reported in other contexts (e.g., native African or American languages) result from superior methodologies or are merely artifacts of benchmark collection. To address this problem, we introduce the FRED Difficulty Metrics, which include the Fertility Ratio (F), Retrieval Proxy (R), Pre-training Exposure (E), and Corpus Diversity (D) and serve as dataset-intrinsic metrics to contextualize reported scores. These metrics reveal that a significant portion of result variability is explained by train-test overlap and pre-training exposure rather than model capability. Additionally, we identify that some languages -- particularly extinct and non-Latin indigenous languages -- suffer from poor tokenization coverage (high token fertility), highlighting a fundamental limitation of transferring models from high-resource languages that lack a shared vocabulary. By providing these indices alongside performance scores, we enable more transparent evaluation of cross-lingual transfer and provide a more reliable foundation for the XLR MT community.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。