让复杂故事生成连贯视频,精准控制角色动作与场景细节。
DreamRunner: Fine-Grained Compositional Story-to-Video Generation with Retrieval-Augmented Motion Adaptation
- 用大模型拆解故事并规划物体布局,实现细粒度控制。
- 通过检索视频自适应运动特征,支持多样化的动作定制。
- 提出时空区域注意力模块,精准绑定角色与动作,提升一致性。
叙事视频生成(SVG)旨在创建符合结构化叙事的多场景连贯视频。现有方法主要依赖大语言模型(LLM)进行高层规划,将故事分解为场景级描述后独立生成并拼接,但在生成复杂单场景视频时表现不佳,因涉及多角色、多事件的协同构图、复杂运动合成及多角色个性化。为此,我们提出DREAMRUNNER:首先利用大语言模型(LLM)对输入脚本进行结构化处理,支持粗粒度场景规划与细粒度物体级布局规划;其次引入检索增强的测试时自适应机制,捕获每场景中目标物体的动作先验,基于检索视频实现多样化动作定制;最后设计空间-时间区域基础3D注意力与先验注入模块SR3AI,实现细粒度物体-动作绑定与逐帧时空语义控制。在多个基准上对比基线,DREAMRUNNER在角色一致性、文本对齐和流畅过渡方面达到领先水平。此外,在T2V-ComBench上的组合式文本到视频生成任务中显著优于基线,验证了其精细条件遵循能力。定性结果进一步表明其生成多对象交互的鲁棒性。
原文摘要 · Abstract (English)
Storytelling video generation (SVG) aims to produce coherent and visually rich multi-scene videos that follow a structured narrative. Existing methods primarily employ LLM for high-level planning to decompose a story into scene-level descriptions, which are then independently generated and stitched together. However, these approaches struggle with generating high-quality videos aligned with the complex single-scene description, as visualizing such complex description involves coherent composition of multiple characters and events, complex motion synthesis and multi-character customization. To address these challenges, we propose DREAMRUNNER, a novel story-to-video generation method: First, we structure the input script using a large language model (LLM) to facilitate both coarse-grained scene planning as well as fine-grained object-level layout planning. Next, DREAMRUNNER presents retrieval-augmented test-time adaptation to capture target motion priors for objects in each scene, supporting diverse motion customization based on retrieved videos, thus facilitating the generation of new videos with complex, scripted motions. Lastly, we propose a novel spatial-temporal region-based 3D attention and prior injection module SR3AI for fine-grained object-motion binding and frame-by-frame spatial-temporal semantic control. We compare DREAMRUNNER with various SVG baselines, demonstrating state-of-the-art performance in character consistency, text alignment, and smooth transitions. Additionally, DREAMRUNNER exhibits strong fine-grained condition-following ability in compositional text-to-video generation, significantly outperforming baselines on T2V-ComBench. Finally, we validate DREAMRUNNER's robust ability to generate multi-object interactions with qualitative examples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。