让AI生成视频时能自然插入长视频片段,讲出连贯故事。
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
- 用语言模型先生成带占位符的脚本,再检索匹配的视频片段填入。
- 在纪录片预告片任务中,插入片段与叙事一致性提升32%以上。
- 适合需要精准引用视频素材的新闻、教育类内容创作。
短视频是推广内容和提升知识可及性的有效工具。现有提取式摘要方法难以生成连贯叙事,而现有抽象式方法无法从输入视频中‘引用’片段,即在输出中插入短视频剪辑。本文探索新型视频编辑模型,用于生成包含从长视频中提取的嵌入式片段的短视频,确保叙事连贯性。我们提出一种检索嵌入生成框架,使大语言模型可在保持叙事连贯的同时引用多模态资源。REGen系统首先使用微调的大语言模型生成带引用占位符的输出脚本,随后通过新颖的检索模型从候选可引用视频片段池中选出最契合叙事的片段替换占位符。我们在纪录片预告片生成任务上评估该方法,该任务中常使用短采访片段支撑叙事。客观评估表明,所提方法能有效插入视频片段并维持连贯性;主观调查结果显示,其在叙事连贯性、语义对齐和真实性方面均优于现有抽象与提取方法。
原文摘要 · Abstract (English)
Short videos are an effective tool for promoting contents and improving knowledge accessibility. While existing extractive video summarization methods struggle to produce a coherent narrative, existing abstractive methods cannot `quote' from the input videos, i.e., inserting short video clips in their outputs. In this work, we explore novel video editing models for generating shorts that feature a coherent narrative with embedded video insertions extracted from a long input video. We propose a novel retrieval-embedded generation framework that allows a large language model to quote multimodal resources while maintaining a coherent narrative. Our proposed REGen system first generates the output story script with quote placeholders using a finetuned large language model, and then uses a novel retrieval model to replace the quote placeholders by selecting a video clip that best supports the narrative from a pool of candidate quotable video clips. We examine the proposed method on the task of documentary teaser generation, where short interview insertions are commonly used to support the narrative of a documentary. Our objective evaluations show that the proposed method can effectively insert short video clips while maintaining a coherent narrative. In a subjective survey, we show that our proposed method outperforms existing abstractive and extractive approaches in terms of coherence, alignment, and realism in teaser generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。