跨领域跨语言的摘要评估新基准,提升评测全面性与质量。
Towards Multi-dimensional Evaluation of LLM Summarization across Domains and Languages
- 构建多维度、多领域中文英文摘要评测框架
- 发现不同模型在跨语言跨领域表现差异显著
- 用多智能体辩论提升人工标注质量,适合评测研究者
文本摘要评估框架在领域覆盖和度量指标上已有所发展,但现有基准仍缺乏领域特异性评估标准,以英语为主,且因推理复杂性导致人工标注困难。为此,我们提出MSumBench,提供中英文双语、多领域的多维摘要评估体系,为各领域设计专用评估标准,并引入多智能体辩论系统以提升标注质量。通过评估八种现代摘要模型,我们发现模型在不同领域和语言间存在明显性能差异。进一步研究大模型作为摘要评价者时,其评价能力与生成能力的相关性,揭示其对自生成摘要存在系统性偏见。该基准数据集已在GitHub公开:https://github.com/DISL-Lab/MSumBench。
原文摘要 · Abstract (English)
Evaluation frameworks for text summarization have evolved in terms of both domain coverage and metrics. However, existing benchmarks still lack domain-specific assessment criteria, remain predominantly English-centric, and face challenges with human annotation due to the complexity of reasoning. To address these, we introduce MSumBench, which provides a multi-dimensional, multi-domain evaluation of summarization in English and Chinese. It also incorporates specialized assessment criteria for each domain and leverages a multi-agent debate system to enhance annotation quality. By evaluating eight modern summarization models, we discover distinct performance patterns across domains and languages. We further examine large language models as summary evaluators, analyzing the correlation between their evaluation and summarization capabilities, and uncovering systematic bias in their assessment of self-generated summaries. Our benchmark dataset is publicly available at https://github.com/DISL-Lab/MSumBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。