让文字生成视频广告更精准,自动控制元素动态布局。
Generating Animated Layouts as Structured Text Representations
- 用分层视觉元素构建结构化文本表示,实现细粒度控制
- 三阶段生成+大模型推理,自动生成带动态轨迹的广告视频
- 适合需要自动化视频广告生成的创作者和营销人员
尽管文本到视频模型取得显著进展,但在视频广告等场景中精确控制文字元素和动画图形仍具挑战。为此,我们提出动画布局生成,通过分层视觉元素的结构化文本表示,扩展静态图形布局以引入时间动态。我们构建了VAKER(Video Ad maKER)——一个结合三阶段生成流程与非结构化文本推理的文本到视频广告生成管线,可无缝集成大语言模型。VAKER通过在特定视频帧中为对象和图形引入动态布局轨迹,实现广告视频的全自动生成。大量实验表明,VAKER在生成视频广告方面显著优于现有方法。
原文摘要 · Abstract (English)
Despite the remarkable progress in text-to-video models, achieving precise control over text elements and animated graphics remains a significant challenge, especially in applications such as video advertisements. To address this limitation, we introduce Animated Layout Generation, a novel approach to extend static graphic layouts with temporal dynamics. We propose a Structured Text Representation for fine-grained video control through hierarchical visual elements. To demonstrate the effectiveness of our approach, we present VAKER (Video Ad maKER), a text-to-video advertisement generation pipeline that combines a three-stage generation process with Unstructured Text Reasoning for seamless integration with LLMs. VAKER fully automates video advertisement generation by incorporating dynamic layout trajectories for objects and graphics across specific video frames. Through extensive evaluations, we demonstrate that VAKER significantly outperforms existing methods in generating video advertisements. Project Page: https://yeonsangshin.github.io/projects/Vaker
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。