arXiv:2504.09071cs.CL2025-04中稿 · the 7th Workshop o…被引 1

小模型写长篇故事摘要时,用计划引导反而没提升,还容易造假。

Exploration of Plan-Guided Summarization for Narrative Texts: the Case of Small Language Models

  • 用叙事结构制定摘要计划,而非只盯日期实体
  • 计划引导的摘要质量和忠实度未优于无计划基线
  • 人类评估发现计划和摘要都常出现幻觉,适合谨慎使用

计划引导的摘要旨在通过将生成内容锚定到源文本,减少小型语言模型(SLMs)的幻觉,通常聚焦于日期或命名实体等细粒度信息。本文研究计划方法在长篇叙事文本摘要任务中的有效性。由于叙事文本长度和复杂性高,忠实摘要难度大。我们分析了现有针对细粒度信息的计划引导方案,并提出一种更高层次的、基于叙事结构的计划形式。结果表明,两种方法均未能显著提升基线模型在摘要质量或忠实度上的表现。人工评估显示,尽管计划引导方法通常与计划一致,但计划本身同样存在幻觉,导致其摘要与无计划模型一样不忠实。本工作为计划引导摘要方法敲响警钟,尤其在长篇复杂文本领域。代码已开源:https://github.com/amazon-science/plan-guided-summarization

原文摘要 · Abstract (English)

Plan-guided summarization attempts to reduce hallucinations in small language models (SLMs) by grounding generated summaries to the source text, typically by targeting fine-grained details such as dates or named entities. In this work, we investigate whether plan-based approaches in SLMs improve summarization in long document, narrative tasks. Narrative texts' length and complexity often mean they are difficult to summarize faithfully. We analyze existing plan-guided solutions targeting fine-grained details, and also propose our own higher-level, narrative-based plan formulation. Our results show that neither approach significantly improves on a baseline without planning in either summary quality or faithfulness. Human evaluation reveals that while plan-guided approaches are often well grounded to their plan, plans are equally likely to contain hallucinations compared to summaries. As a result, the plan-guided summaries are just as unfaithful as those from models without planning. Our work serves as a cautionary tale to plan-guided approaches to summarization, especially for long, complex domains such as narrative texts. Code available at https://github.com/amazon-science/plan-guided-summarization

摘要生成小模型幻觉抑制叙事文本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。