arXiv:2509.17667cs.CL2025-09被引 1

为13种印度语言构建翻译评估数据集并训练新模型

Crosslingual Optimized Metric for Translation Assessment of Indian Languages

  • 基于13种印度语言的21个方向,构建大规模人工评分数据集
  • 新模型COMTAIL在含印度语言的翻译对上显著优于现有方法
  • 支持多领域、多语言组别敏感性分析,适合本地化评估研究

自动翻译评估因语言间拼写、形态、句法和语义的丰富性与差异性而面临挑战。传统基于字符串的度量如BLEU虽广泛应用,但其局限性日益显现。尽管学习型神经度量缓解了部分问题,但在多数语言(尤其是非主流高资源语言对)中仍受限于高质量标注数据稀缺。本文针对这些不足,构建了涵盖13种印度语言、21个翻译方向的大规模人工评分数据集,并在此基础上训练名为跨语言优化翻译评估度量(COMTAIL)的神经评估模型。最优模型在至少包含一种印度语言的翻译对评估中显著超越现有最先进方法。通过一系列消融实验,验证了该度量对领域、翻译质量及语言分组变化的敏感性。本文同时公开了COMTAIL数据集与配套模型。

原文摘要 · Abstract (English)

Automatic evaluation of translation remains a challenging task owing to the orthographic, morphological, syntactic and semantic richness and divergence observed across languages. String-based metrics such as BLEU have previously been extensively used for automatic evaluation tasks, but their limitations are now increasingly recognized. Although learned neural metrics have helped mitigate some of the limitations of string-based approaches, they remain constrained by a paucity of gold evaluation data in most languages beyond the usual high-resource pairs. In this present work we address some of these gaps. We create a large human evaluation ratings dataset for 13 Indian languages covering 21 translation directions and then train a neural translation evaluation metric named Cross-lingual Optimized Metric for Translation Assessment of Indian Languages (COMTAIL) on this dataset. The best performing metric variants show significant performance gains over previous state-of-the-art when adjudging translation pairs with at least one Indian language. Furthermore, we conduct a series of ablation studies to highlight the sensitivities of such a metric to changes in domain, translation quality, and language groupings. We release both the COMTAIL dataset and the accompanying metric models.

翻译评估多语言印度语言神经度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。