arXiv:2410.05586cs.CVcs.AI2024-10ICLR被引 5

用长纪录片生成吸引人的预告片,解决音画对齐与事实准确难题。

TeaserGen: Generating Teasers for Long Documentaries

  • 分两阶段生成:先用大模型写旁白,再选匹配画面
  • 在1269部纪录片上验证,预训练视觉模型更准
  • 适合影视推广、内容创作与多模态生成研究者

预告片是娱乐、商业和教育领域推广内容的有效工具。然而,为长视频制作有效预告片极具挑战性,需对输入视频进行长程多模态建模,同时保持音画对齐、管理场景切换并确保输出内容的事实准确性。由于缺乏公开数据集,该方向进展受限。本文提出DocumentaryNet,一个包含1,269部纪录片及其对应预告片的多模态数据集,涵盖视频、语音、音乐、音效和旁白等数据流。基于此,我们设计TeaserGen系统:第一阶段使用预训练大语言模型从纪录片旁白生成预告片旁白;第二阶段通过语言-视觉模型选择最相关的视觉内容。针对旁白与画面匹配,我们探索两种方法:基于预训练对比语言-视觉模型的方法,以及学习映射关系的深度序列模型。实验表明,预训练方法在识别相关视觉内容方面优于直接训练的自回归模型。

原文摘要 · Abstract (English)

Teasers are an effective tool for promoting content in entertainment, commercial and educational fields. However, creating an effective teaser for long videos is challenging for it requires long-range multimodal modeling on the input videos, while necessitating maintaining audiovisual alignments, managing scene changes and preserving factual accuracy for the output teasers. Due to the lack of a publicly-available dataset, progress along this research direction has been hindered. In this work, we present DocumentaryNet, a collection of 1,269 documentaries paired with their teasers, featuring multimodal data streams of video, speech, music, sound effects and narrations. With DocumentaryNet, we propose a new two-stage system for generating teasers from long documentaries. The proposed TeaserGen system first generates the teaser narration from the transcribed narration of the documentary using a pretrained large language model, and then selects the most relevant visual content to accompany the generated narration through language-vision models. For narration-video matching, we explore two approaches: a pretraining-based model using pretrained contrastive language-vision models and a deep sequential model that learns the mapping between the narrations and visuals. Our experimental results show that the pretraining-based approach is more effective at identifying relevant visual content than directly trained deep autoregressive models.

视频生成多模态预告片语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。