让AI生成连贯长视频,支持多模态控制与自动叙事规划。
Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal Conditioning
- 用大语言模型做故事规划,动态记忆库防画面漂移。
- 在电商广告等场景下实现跨镜头视觉一致性,效果领先。
- 自建33万张带叙事标注的电商视频数据集,填补评估空白。
我们提出「Narrative Weaver」,一种解决生成式AI中多模态可控、长序列、视觉一致性难题的新框架。现有模型虽能生成高质量短时视觉内容,却难以保持长序列中的叙事连贯性与视觉一致性,制约了影视制作和电商广告等实际应用。Narrative Weaver首次融合细粒度控制、自动叙事规划与长程一致性三大能力:结合多模态大语言模型(MLLM)进行高层叙事规划,并引入动态记忆库模块的细粒度控制机制,有效防止视觉漂移。为支持实用部署,设计渐进式多阶段训练策略,高效利用预训练模型,在有限数据下达到顶尖性能。鉴于缺乏合适评估基准,我们构建并发布首个综合性数据集——电商广告视频分镜数据集(E-commerce Advertising Video Storyboard Dataset, EAVSD),包含超过33万张高质量图像及丰富的叙事标注。在三种不同场景(可控多场景生成、自主讲故事、电商广告)的大量实验中,验证了该方法的优越性,为人工智能驱动的内容创作开辟新可能。
原文摘要 · Abstract (English)
We present "Narrative Weaver", a novel framework that addresses a fundamental challenge in generative AI: achieving multi-modal controllable, long-range, and consistent visual content generation. While existing models excel at generating high-fidelity short-form visual content, they struggle to maintain narrative coherence and visual consistency across extended sequences - a critical limitation for real-world applications such as filmmaking and e-commerce advertising. Narrative Weaver introduces the first holistic solution that seamlessly integrates three essential capabilities: fine-grained control, automatic narrative planning, and long-range coherence. Our architecture combines a Multimodal Large Language Model (MLLM) for high-level narrative planning with a novel fine-grained control module featuring a dynamic Memory Bank that prevents visual drift. To enable practical deployment, we develop a progressive, multi-stage training strategy that efficiently leverages existing pre-trained models, achieving state-of-the-art performance even with limited training data. Recognizing the absence of suitable evaluation benchmarks, we construct and release the E-commerce Advertising Video Storyboard Dataset (EAVSD) - the first comprehensive dataset for this task, containing over 330K high-quality images with rich narrative annotations. Through extensive experiments across three distinct scenarios (controllable multi-scene generation, autonomous storytelling, and e-commerce advertising), we demonstrate our method's superiority while opening new possibilities for AI-driven content creation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。