arXiv:2507.02790cs.CVcs.CL2025-07EMNLP被引 7

用多模态理解让机器自动剪出更连贯的短视频

From Long Videos to Engaging Clips: A Human-Inspired Video Editing Framework with Multimodal Narrative Understanding

  • 结合角色、对话和叙事分析,实现视频内容的全面理解
  • 在800段剧集和500条广告上显著提升剪辑连贯性
  • 适合想做智能视频剪辑或内容生成的研究者

在线视频内容的爆炸式增长,尤其在短视频平台,催生了将长视频高效压缩为精炼吸引人片段的需求。现有自动编辑方法主要依赖ASR转录文本和端到端片段选择,常忽略丰富视觉信息,导致输出不连贯。本文提出一种受人类启发的自动视频编辑框架HIVE,通过多模态大语言模型实现角色提取、对话分析与叙事摘要,获得对视频内容的整体理解。为增强连贯性,引入场景级分割,并将编辑过程分解为三个子任务:亮点检测、开头结尾选择、无关内容剔除。为推动该领域研究,我们构建了DramaAD数据集,包含超过800段短剧集和500条专业剪辑广告片段。实验表明,该框架在通用与广告导向编辑任务中均持续优于现有基线,显著缩小了自动剪辑与人工剪辑的质量差距。

原文摘要 · Abstract (English)

The rapid growth of online video content, especially on short video platforms, has created a growing demand for efficient video editing techniques that can condense long-form videos into concise and engaging clips. Existing automatic editing methods predominantly rely on textual cues from ASR transcripts and end-to-end segment selection, often neglecting the rich visual context and leading to incoherent outputs. In this paper, we propose a human-inspired automatic video editing framework (HIVE) that leverages multimodal narrative understanding to address these limitations. Our approach incorporates character extraction, dialogue analysis, and narrative summarization through multimodal large language models, enabling a holistic understanding of the video content. To further enhance coherence, we apply scene-level segmentation and decompose the editing process into three subtasks: highlight detection, opening/ending selection, and pruning of irrelevant content. To facilitate research in this area, we introduce DramaAD, a novel benchmark dataset comprising over 800 short drama episodes and 500 professionally edited advertisement clips. Experimental results demonstrate that our framework consistently outperforms existing baselines across both general and advertisement-oriented editing tasks, significantly narrowing the quality gap between automatic and human-edited videos.

视频剪辑多模态叙事理解自动生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。