arXiv:2410.14666cs.CLcs.AI2024-10NAACL被引 3

用角色关联图增强电影剧本摘要,更懂人物关系与情节脉络

DiscoGraMS: Enhancing Movie Screen-Play Summarization using Movie Character-Aware Discourse Graph

  • 构建角色感知的对话图,捕捉人物互动与隐含关系
  • 相比传统模型,能更好保留长程依赖与关键情节信息
  • 适合做剧本摘要、问答系统或角色重要性分析的研究者

电影剧本摘要相较于常规文档摘要面临独特挑战:剧本不仅篇幅长,且包含复杂的人物、对白与场景交互,存在大量直接与间接关系及语境细节,难以被机器学习模型准确捕捉。现有方法多基于微调Transformer预训练模型,但常无法有效建模长距离依赖和潜在关系,易出现“中间内容丢失”问题。为此,我们提出DiscoGraMS,将电影剧本表示为角色感知的对话图(CaD Graph),该结构适用于摘要、问答与显著性检测等下游任务。模型旨在完整保留关键信息,实现更全面、忠实的剧本表征。我们还设计了一种基于后期融合的基线方法,将图结构与文本内容结合,初步结果展现出良好前景。

原文摘要 · Abstract (English)

Summarizing movie screenplays presents a unique set of challenges compared to standard document summarization. Screenplays are not only lengthy, but also feature a complex interplay of characters, dialogues, and scenes, with numerous direct and subtle relationships and contextual nuances that are difficult for machine learning models to accurately capture and comprehend. Recent attempts at screenplay summarization focus on fine-tuning transformer-based pre-trained models, but these models often fall short in capturing long-term dependencies and latent relationships, and frequently encounter the "lost in the middle" issue. To address these challenges, we introduce DiscoGraMS, a novel resource that represents movie scripts as a movie character-aware discourse graph (CaD Graph). This approach is well-suited for various downstream tasks, such as summarization, question-answering, and salience detection. The model aims to preserve all salient information, offering a more comprehensive and faithful representation of the screenplay's content. We further explore a baseline method that combines the CaD Graph with the corresponding movie script through a late fusion of graph and text modalities, and we present very initial promising results.

剧本摘要图神经网络角色关系多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。