arXiv:2605.03244cs.CLcs.AI2026-05KDD

用叙事理论提炼电影剧本核心情节,实现更精准的自动摘要

S^2tory: Story Spine Distillation for Movie Script Summarization

论文配图:S^2tory: Story Spine Distillation for Movie Script Summarization
图 1 · 摘自论文原文
  • 基于角色发展轨迹识别推动剧情的核心事件
  • 在MovieSum上实现约3.5倍压缩率且语义保真度领先
  • 适合需要理解复杂非线性叙事的研究者与开发者

电影剧本因非线性、交叉叙述结构,使传统表面显著性方法难以保留核心剧情进展。为此,我们提出S^2tory(故事骨架蒸馏)框架,基于叙事学理论,利用角色发展轨迹识别驱动叙事前进的“剧情核”事件,过滤仅增强氛围或情感的附属事件。叙事专家代理(NEAgent)执行受理论约束的推理,其知识蒸馏后指导小型模型识别剧情核;另一模型则基于这些核心事件生成摘要。在MovieSum数据集上的实验表明,该方法实现约3.5倍压缩率下的最优语义保真度;在BookSum上的零样本评估验证了出色的跨领域泛化能力。人工评估进一步证明,叙事学理论是建模复杂非线性叙事不可或缺的基础。

原文摘要 · Abstract (English)

Movie scripts pose a fundamental challenge for automatic summarization due to their non-linear, cross-cut narrative structure, which makes surface-level saliency methods ineffective at preserving core story progression. To address this, we introduce S^2tory (Story Spine Distillation), a narratology-grounded framework that leverages character development trajectories to identify plot nuclei, the essential events that drive the narrative forward, while filtering out peripheral satellite events that merely enrich atmosphere or emotion. Our Narrative Expert Agent (NEAgent) performs theory-constrained reasoning, whose distilled knowledge conditions a small model to identify plot nuclei. Another model then uses these plot nuclei to generate the summary. Experiments on the MovieSum dataset demonstrate state-of-the-art semantic fidelity at approximately 3.5x compression, and zero-shot evaluation on BookSum confirms strong out-of-domain generalization. Human evaluation further validates that narratological theory provides an indispensable foundation for modeling complex, non-linear narratives.

剧本摘要叙事学知识蒸馏

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。