仅用视频自动生成剧本和剧情摘要,突破依赖文本的局限
ScreenWriter: Automatic Screenplay Generation and Movie Summarisation
- 基于视觉向量序列分割视频场景,结合演员人脸库识别角色名称
- 生成含对白、角色名、场景切换和视觉描述的完整剧本
- 可直接从视频生成高质量剧情摘要,适合影视内容分析场景
创意视频内容激增催生了自动文本描述需求,以便用户快速回顾关键情节。然而,电影数量庞大且更新迅速,自动摘要面临识别角色意图和长时序依赖的挑战。现有方法严重依赖剧本文本输入,适用性受限。本文提出自动剧本生成任务及ScreenWriter方法,仅使用视频即可输出包含对白、角色名、场景分隔和视觉描述的剧本。该方法引入新算法,基于视觉向量序列划分场景,并利用演员人脸数据库解决角色命名难题。进一步采用基于场景分隔的分层摘要法生成剧情概要。在扩展后的MovieSum数据集上测试,结果优于多个依赖真实剧本的对比模型。
原文摘要 · Abstract (English)
The proliferation of creative video content has driven demand for textual descriptions or summaries that allow users to recall key plot points or get an overview without watching. The volume of movie content and speed of turnover motivates automatic summarisation, which is nevertheless challenging, requiring identifying character intentions and very long-range temporal dependencies. The few existing methods attempting this task rely heavily on textual screenplays as input, greatly limiting their applicability. In this work, we propose the task of automatic screenplay generation, and a method, ScreenWriter, that operates only on video and produces output which includes dialogue, speaker names, scene breaks, and visual descriptions. ScreenWriter introduces a novel algorithm to segment the video into scenes based on the sequence of visual vectors, and a novel method for the challenging problem of determining character names, based on a database of actors' faces. We further demonstrate how these automatic screenplays can be used to generate plot synopses with a hierarchical summarisation method based on scene breaks. We test the quality of the final summaries on the recent MovieSum dataset, which we augment with videos, and show that they are superior to a number of comparison models which assume access to goldstandard screenplays.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。