对比大模型与人工摘要,发现人类仍更准确可靠
Summarization is Not Dead Yet

- 用多维度评测方法比较人类与大模型摘要质量
- 人类摘要在信息量和事实准确性上明显领先
- 适合关注摘要可信度与真实性的研究者阅读
大型语言模型(LLMs)的发展催生了模型生成摘要已超越或媲美人工参考的论断,引发对摘要是否仍是开放研究问题的质疑。我们通过涵盖多种数据集和先进大模型的多轨道评估重新审视这一观点,结合受控的人类评估、去偏的LLM作为裁判协议、外部知识的事实性验证以及语料级语言学分析。结果表明,人类参考在信息量和忠实度方面仍具优势,而大模型输出主要在表面连贯性和流畅性上更受青睐。事实性验证显示,人类摘要在涉及推理或综合的陈述上更为可靠,语言学分析揭示不同模型间存在风格同质化现象。这些观察表明,当前大模型虽提升了摘要质量的下限,但其上限仍低于人类能力。
原文摘要 · Abstract (English)
The progress of large language models (LLMs) has fueled claims that model-generated summaries rival or even surpass human-written references, raising questions about whether summarization remains an open research problem. We re-examine this narrative through a multi-track evaluation covering diverse datasets and state-of-the-art LLMs, combining controlled human assessment, bias-mitigated LLM-as-Judge protocols, factuality verification against external knowledge, and corpus-level linguistic analysis. Our findings reveal a more nuanced landscape in which human references continue to demonstrate advantages in informativeness and faithfulness, whereas LLM outputs are preferred mainly for surface-level coherence and fluency. Factuality verification indicates that human references remain more reliable, particularly for claims involving reasoning or synthesis, and linguistic analysis uncovers a pattern of stylistic homogeneity across different models. These observations suggest that current LLMs have raised the floor of summarization quality, but the ceiling of their performance remains below human capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。