arXiv:2508.08487cs.CVcs.AI2025-08Conference of the …被引 19

MAViS用多智能体协作生成高质量长视频故事,支持从想法到成片的全流程创作。

MAViS: A Multi-Agent Framework for Long-Sequence Video Storytelling

  • 设计多阶段智能体协同框架,每阶段遵循探索-审视-优化原则保证输出完整
  • 在脚本与生成工具间引入写作指南,提升生成内容兼容性与质量
  • 首次实现视频+叙事+背景音乐的多模态输出,适合创意工作者快速试错

尽管近期取得进展,长序列视频生成仍面临辅助能力弱、视觉质量不佳和表现力有限等挑战。为此,我们提出MAViS,一种多智能体协作框架,通过高效将创意转化为视觉叙事,助力长序列视频故事创作。MAViS在脚本撰写、镜头设计、角色建模、关键帧生成、视频动画及音频生成等多个阶段协调专用智能体,各阶段均遵循‘探索-审视-增强’(3E)原则,确保中间结果完整性。针对当前生成模型能力限制,提出脚本编写指南以优化脚本与生成工具的适配性。实验表明,MAViS在辅助能力、视觉质量和视频表现力上达到当前最优水平。其模块化架构可灵活集成多种生成模型与工具。仅需简要想法描述,即可快速生成高质量、完整的长序列视频,探索多样化的视觉叙事方向。据我们所知,MAViS是首个提供多模态设计输出——含叙事与背景音乐的视频——的框架。

原文摘要 · Abstract (English)

Despite recent advances, long-sequence video generation frameworks still suffer from significant limitations: poor assistive capability, suboptimal visual quality, and limited expressiveness. To mitigate these limitations, we propose MAViS, a multi-agent collaborative framework designed to assist in long-sequence video storytelling by efficiently translating ideas into visual narratives. MAViS orchestrates specialized agents across multiple stages, including script writing, shot designing, character modeling, keyframe generation, video animation, and audio generation. In each stage, agents operate under the 3E Principle -- Explore, Examine, and Enhance -- to ensure the completeness of intermediate outputs. Considering the capability limitations of current generative models, we propose the Script Writing Guidelines to optimize compatibility between scripts and generative tools. Experimental results demonstrate that MAViS achieves state-of-the-art performance in assistive capability, visual quality, and video expressiveness. Its modular framework further enables scalability with diverse generative models and tools. With just a brief idea description, MAViS enables users to rapidly explore diverse visual storytelling and creative directions for sequential video generation by efficiently producing high-quality, complete long-sequence videos. To the best of our knowledge, MAViS is the only framework that provides multimodal design output -- videos with narratives and background music.

视频生成多智能体长视频叙事生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。