arXiv:2505.12983cs.CLcs.AI2025-05ACL被引 1

评测大模型多语言摘要能力,发现微调后表现超预期但事实性问题仍存。

An Empirical Study of Many-to-Many Summarization with Large Language Models

  • 重构5大领域6语言共47.8万样本数据集,支持多语言摘要研究。
  • 微调后开源大模型在自动评估中超越零样本模型(含GPT-4)。
  • 指令微调提升性能但可能加剧事实错误,需关注真实性控制。

多对多摘要(M2MS)旨在处理任意语言的文档并生成对应语言的摘要。近期大语言模型(LLMs)展现出强大的多语言能力,具备实现真实应用中M2MS的潜力。本文开展系统性实证研究:首先基于8个既有领域数据集重构数据,形成包含47.8万样本、覆盖5个领域和6种语言的新数据集;随后在零样本与指令微调两种方式下基准测试18个LLMs,同时对比微调的传统模型(如mBART)。实验表明,零样本LLMs表现已接近微调传统模型;指令微调后,开源LLMs显著提升多语言摘要能力,且在自动评估中优于零样本模型(包括GPT-4)。此外,任务优化未损害其通用任务求解能力。然而,人工评估揭示LLMs仍存在事实性错误,且指令微调可能加剧该问题。因此,在实际构建大模型摘要系统时,如何控制事实错误成为关键挑战,值得未来深入研究。

原文摘要 · Abstract (English)

Many-to-many summarization (M2MS) aims to process documents in any language and generate the corresponding summaries also in any language. Recently, large language models (LLMs) have shown strong multi-lingual abilities, giving them the potential to perform M2MS in real applications. This work presents a systematic empirical study on LLMs' M2MS ability. Specifically, we first reorganize M2MS data based on eight previous domain-specific datasets. The reorganized data contains 47.8K samples spanning five domains and six languages, which could be used to train and evaluate LLMs. Then, we benchmark 18 LLMs in a zero-shot manner and an instruction-tuning manner. Fine-tuned traditional models (e.g., mBART) are also conducted for comparisons. Our experiments reveal that, zero-shot LLMs achieve competitive results with fine-tuned traditional models. After instruct-tuning, open-source LLMs can significantly improve their M2MS ability, and outperform zero-shot LLMs (including GPT-4) in terms of automatic evaluations. In addition, we demonstrate that this task-specific improvement does not sacrifice the LLMs' general task-solving abilities. However, as revealed by our human evaluation, LLMs still face the factuality issue, and the instruction tuning might intensify the issue. Thus, how to control factual errors becomes the key when building LLM summarizers in real applications, and is worth noting in future research.

多语言摘要生成大模型事实性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。