让文字描述自动匹配并拼接视频片段,生成连贯短片。
Text-Video Multi-Grained Integration for Video Moment Montage
- 融合文本与视频的镜头级和帧级特征进行多粒度对齐
- 在自建数据集上显著提升片段定位与拼接准确率
- 适合视频自动化剪辑、智能内容生成等场景
在线短视频平台的兴起推动了用户对短视频编辑的需求。然而,手动选择、裁剪并拼接原始画面以生成连贯高质量视频仍耗时费力。为加速该过程,我们提出一种新任务——视频片段蒙太奇(Video Moment Montage, VMM),旨在根据预设叙述文本精准定位对应视频片段,并将其有序排列生成符合描述的完整视频。挑战在于需提取精确的时间段,并保证句内与句间语义一致性,因一句文本可能需要多个视频片段的裁剪与拼接。为此,我们提出一种新颖的文本-视频多粒度融合方法(TV-MGI),高效融合脚本文本特征与视频的镜头级和帧级特征,实现文本与视频内容在全局与细粒度层面的对齐。为促进该领域研究,我们构建了大规模专用数据集多重句子带镜头数据集(MSSD)。在MSSD上的大量实验表明,该框架优于基线方法。
原文摘要 · Abstract (English)
The proliferation of online short video platforms has driven a surge in user demand for short video editing. However, manually selecting, cropping, and assembling raw footage into a coherent, high-quality video remains laborious and time-consuming. To accelerate this process, we focus on a user-friendly new task called Video Moment Montage (VMM), which aims to accurately locate the corresponding video segments based on a pre-provided narration text and then arrange these video clips to create a complete video that aligns with the corresponding descriptions. The challenge lies in extracting precise temporal segments while ensuring intra-sentence and inter-sentence context consistency, as a single script sentence may require trimming and assembling multiple video clips. To address this problem, we present a novel \textit{Text-Video Multi-Grained Integration} method (TV-MGI) that efficiently fuses text features from the script with both shot-level and frame-level video features, which enables the global and fine-grained alignment between the video content and the corresponding textual descriptions in the script. To facilitate further research in this area, we introduce the Multiple Sentences with Shots Dataset (MSSD), a large-scale dataset designed explicitly for the VMM task. We conduct extensive experiments on the MSSD dataset to demonstrate the effectiveness of our framework compared to baseline methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。