arXiv:2503.17039cs.CLcs.AI2025-03被引 6

对比中英文摘要评价指标,发现大模型裁判在西语巴斯克语上表现差异显著。

Summarization Metrics for Spanish and Basque: Do Automatic Scores and LLM-Judges Correlate with Humans?

  • 构建双语摘要评估数据集BASSE,含2040条人工标注摘要。
  • 商用大模型裁判相关性最高,开源模型表现差,传统指标次之。
  • 适合多语言摘要评估、跨语言大模型研究者参考。

现有自动摘要评价指标与大模型裁判研究主要集中于英语,限制了对其他语言的评估效果理解。本文通过新构建的BASSE数据集(巴斯克语和西班牙语摘要评估),收集了2040条抽象式摘要的人工判断数据,这些摘要由人工或五种大模型在四种提示下生成。每条摘要由标注者按5分制对连贯性、一致性、流畅性、相关性和5W1H五个维度评分。我们基于此数据重新评估了传统自动指标及在英语任务中表现优异的LLM-as-a-Judge模型。结果表明,当前专有大模型裁判与人工判断相关性最高,其次是针对特定标准设计的自动指标,而开源大模型裁判表现较差。我们公开发布BASSE数据集及代码,并提供首个大规模巴斯克语摘要数据集,包含22,525篇新闻文章及其标题。

原文摘要 · Abstract (English)

Studies on evaluation metrics and LLM-as-a-Judge models for automatic text summarization have largely been focused on English, limiting our understanding of their effectiveness in other languages. Through our new dataset BASSE (BAsque and Spanish Summarization Evaluation), we address this situation by collecting human judgments on 2,040 abstractive summaries in Basque and Spanish, generated either manually or by five LLMs with four different prompts. For each summary, annotators evaluated five criteria on a 5-point Likert scale: coherence, consistency, fluency, relevance, and 5W1H. We use these data to reevaluate traditional automatic metrics used for evaluating summaries, as well as several LLM-as-a-Judge models that show strong performance on this task in English. Our results show that currently proprietary judge LLMs have the highest correlation with human judgments, followed by criteria-specific automatic metrics, while open-sourced judge LLMs perform poorly. We release BASSE and our code publicly, along with the first large-scale Basque summarization dataset containing 22,525 news articles with their subheads.

摘要评估多语言大模型裁判巴斯克语

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。