通过对比人类与大模型的小说摘要,揭示其概念关注差异。
Attention Flows: Tracing LLM Conceptual Engagement via Story Summaries

- 用150个真人摘要对齐章节,量化总结中的概念聚焦点。
- 发现大模型更关注文本结尾,而人类更均衡分布注意力。
- 结果可解释长文理解退化,适合研究模型推理机制者参考。
尽管大语言模型的上下文长度持续增长,但其在长篇文本中整合信息的能力并未同步提升。我们评估了一项关键理解任务:小说摘要生成。人类撰写摘要时会凸显其认为重要的情节,因此通过对比人类与大模型生成的摘要,可判断模型是否模拟了人类的概念关注模式。我们首先将150份真人撰写的小说摘要中的句子与对应章节对齐,验证了该对齐任务的难度,反映出摘要生成的复杂性。随后,我们为每篇参考文本生成并对齐九个顶尖大模型的摘要。对比结果显示,人类与模型在风格和叙事关注分布上存在差异,模型更倾向关注文本末尾内容。这一发现有助于解释大模型在长文理解上的性能下降,并为未来改进提供方向。数据集已公开,支持后续研究。
原文摘要 · Abstract (English)
Although LLM context lengths have grown, there is evidence that their ability to integrate information across long-form texts has not kept pace. We evaluate one such understanding task: generating summaries of novels. When human authors of summaries compress a story, they reveal what they consider narratively important. Therefore, by comparing human and LLM-authored summaries, we can assess whether models mirror human patterns of conceptual engagement with texts. To measure conceptual engagement, we align sentences from 150 human-written novel summaries with the specific chapters they reference. We demonstrate the difficulty of this alignment task, which indicates the complexity of summarization as a task. We then generate and align additional summaries by nine state-of-the-art LLMs for each of the 150 reference texts. Comparing the human and model-authored summaries, we find both stylistic differences between the texts and differences in how humans and LLMs distribute their focus throughout a narrative, with models emphasizing the ends of texts. Comparing human narrative engagement with model attention mechanisms suggests explanations for degraded narrative comprehension and targets for future development. We release our dataset to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。