arXiv:2507.08342cs.CL2025-07ACL被引 7

跨语言摘要评估发现:融合语用n-gram指标效果差,需用神经网络评估模型

Beyond N-Grams: Rethinking Evaluation Metrics and Strategies for Multilingual Abstractive Summarization

  • 对比8种语言的n-gram与神经评估方法,发现融合语中n-gram相关性更低
  • 正确分词可显著提升融合语的评估相关性,甚至逆转负面趋势
  • 专门训练的神经评估模型(如COMET)在低资源语言中表现更优

自动摘要评估广泛使用基于n-gram的度量(如ROUGE),尽管其对英语仍具参考价值,但在其他语言中的适用性尚不明确。本文系统评估了多语言生成任务中各类评估指标的有效性,涵盖来自四种类型语言家族(黏着型、孤立型、低融合型、高融合型)的八种语言,覆盖低/高资源场景。结果表明,评估指标对语言类型敏感:在融合型语言中,n-gram指标与人工判断的相关性低于孤立型和黏着型语言。通过恰当分词可显著缓解此问题,甚至逆转负向趋势。此外,专为评估训练的神经指标(如COMET)始终优于其他神经指标,且在低资源语言中与人工判断相关性更高。研究揭示了n-gram指标在融合语言中的局限性,呼吁加强针对评估任务的神经模型投入。

原文摘要 · Abstract (English)

Automatic n-gram based metrics such as ROUGE are widely used for evaluating generative tasks such as summarization. While these metrics are considered indicative (even if imperfect) of human evaluation for English, their suitability for other languages remains unclear. To address this, we systematically assess evaluation metrics for generation both n-gram-based and neural based to evaluate their effectiveness across languages and tasks. Specifically, we design a large-scale evaluation suite across eight languages from four typological families: agglutinative, isolating, low-fusional, and high-fusional, spanning both low- and high-resource settings, to analyze their correlation with human judgments. Our findings highlight the sensitivity of evaluation metrics to the language type. For example, in fusional languages, n-gram-based metrics show lower correlation with human assessments compared to isolating and agglutinative languages. We also demonstrate that proper tokenization can significantly mitigate this issue for morphologically rich fusional languages, sometimes even reversing negative trends. Additionally, we show that neural-based metrics specifically trained for evaluation, such as COMET, consistently outperform other neural metrics and better correlate with human judgments in low-resource languages. Overall, our analysis highlights the limitations of n-gram metrics for fusional languages and advocates for greater investment in neural-based metrics trained for evaluation tasks.

摘要生成多语言评估指标神经模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。