针对印度语机器翻译与摘要评估,提出新基准并发现主流指标表现差异。
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
- 构建涵盖6种印度语的29个评估指标大样本测试集。
- 大模型评估器在段落和系统级与人工判断最匹配。
- 指标对异常值敏感,且在内容忠实度与流畅性上各有侧重。
尽管自动评估指标推动了机器翻译(MT)和文本摘要(TS)的发展,但现有指标几乎仅针对英语等高资源语言开发和验证。这一局限使超过15亿人使用的印度语言被严重忽视,引发当前评估方法普适性的质疑。为此,我们提出ITEM——一个大规模基准,系统评估29种自动指标在六种主要印度语言中与人工判断的一致性,并引入细粒度标注。全面评估涵盖与人工判断的一致性、对异常值的敏感性、语言特异性可靠性、指标间相关性及受控扰动下的鲁棒性,得出四项核心发现:(1)基于大模型的评估器在段落和系统层级与人工判断一致性最强;(2)异常值显著影响指标与人工判断的一致性;(3)在摘要任务中,指标更擅长捕捉内容忠实度;而在翻译任务中,更准确反映流畅性;(4)不同指标在面对多种扰动时表现出不同的鲁棒性与敏感性。这些发现为印度语言中评估指标的设计与优化提供了关键指导。
原文摘要 · Abstract (English)
While automatic metrics drive progress in Machine Translation (MT) and Text Summarization (TS), existing metrics have been developed and validated almost exclusively for English and other high-resource languages. This narrow focus leaves Indian languages, spoken by over 1.5 billion people, largely overlooked, casting doubt on the universality of current evaluation practices. To address this gap, we introduce ITEM, a large-scale benchmark that systematically evaluates the alignment of 29 automatic metrics with human judgments across six major Indian languages, enriched with fine-grained annotations. Our extensive evaluation, covering agreement with human judgments, sensitivity to outliers, language-specific reliability, inter-metric correlations, and resilience to controlled perturbations reveals four central findings: (1) LLM-based evaluators show the strongest alignment with human judgments at both segment and system levels; (2) outliers exert a significant impact on metric-human agreement; (3) In TS, metrics are more effective at capturing content fidelity, whereas in MT, they better reflect fluency; and (4) Metrics differ in their robustness and sensitivity when subjected to diverse perturbations. Collectively, these findings offer critical guidance for advancing metric design and evaluation in Indian languages.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。