arXiv:2501.05884cs.CV2025-01被引 7

用文字直接控制视频剪辑,让广告制作更智能高效。

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs

  • 通过多模态大模型实现文本到视频编辑的端到端控制。
  • 采用帧率增强与快慢处理结合,提升时空信息理解能力。
  • 适合需要快速生成定制化短视频的广告与内容创作者。

短视频内容的爆发式增长催生了对高效自动化视频编辑解决方案的需求,核心挑战在于理解视频内容并根据用户需求精准调整。为此,我们提出一种创新的端到端基础框架,实现了对最终视频内容编辑的精确控制。利用多模态大语言模型(MLLMs)的灵活性与泛化能力,我们定义了清晰的输入-输出映射以实现高效视频生成。为增强模型对视频内容的处理与理解能力,引入密集帧率与快慢处理相结合的技术策略,显著提升了对时间与空间信息的提取与理解效果。此外,我们设计了文本到编辑机制,用户仅需输入文本即可达成期望的视频结果,从而大幅提升编辑质量与可控性。在广告数据集上的综合实验表明,该方法不仅表现出显著有效性,且在公开数据集上也得出具有普遍适用性的结论。

原文摘要 · Abstract (English)

The exponential growth of short-video content has ignited a surge in the necessity for efficient, automated solutions to video editing, with challenges arising from the need to understand videos and tailor the editing according to user requirements. Addressing this need, we propose an innovative end-to-end foundational framework, ultimately actualizing precise control over the final video content editing. Leveraging the flexibility and generalizability of Multimodal Large Language Models (MLLMs), we defined clear input-output mappings for efficient video creation. To bolster the model's capability in processing and comprehending video content, we introduce a strategic combination of a denser frame rate and a slow-fast processing technique, significantly enhancing the extraction and understanding of both temporal and spatial video information. Furthermore, we introduce a text-to-edit mechanism that allows users to achieve desired video outcomes through textual input, thereby enhancing the quality and controllability of the edited videos. Through comprehensive experimentation, our method has not only showcased significant effectiveness within advertising datasets, but also yields universally applicable conclusions on public datasets.

视频生成多模态大模型文本控制广告制作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。