ViMax通过多智能体协作实现长视频连贯生成
ViMax: Agentic Video Generation

- 多智能体分工协作,统一规划剧情与视觉一致性
- 支持跨场景角色与环境状态持续追踪
- 适合需要长序列叙事的视频生成任务
长视频生成需要系统性叙事规划和视觉一致性,现有短片段方法无法满足。当前方法生成孤立片段,缺乏叙事结构,且难以维持角色与环境的一致性。我们提出ViMax,一种基于多智能体协同的视频生成框架,通过专业化组件协商叙事决策、视觉连续性与制作质量。框架采用分层叙事引擎结合检索增强生成以保证全局故事连贯性,并引入依赖感知的视觉一致性机制,跨时间边界追踪角色与环境状态;同时,基于视觉语言模型(VLM)的智能体持续监控并优化叙事连贯性与视觉保真度。该框架实现了智能体间的协同合作,生成具有完整叙事结构与跨多场景视觉一致性的长篇内容。
原文摘要 · Abstract (English)
Long-form video generation requires systematic narrative planning and visual consistency that current short-clip methods cannot provide. Existing methods generate isolated sequences without narrative structure and lack mechanisms for maintaining character and environmental consistency across scenes. We present ViMax, an agentic video generation framework that addresses video creation through coordinated multi-agent collaboration where specialized components negotiate narrative decisions, visual continuity, and production quality. Our framework employs a hierarchical narrative engine with retrieval-augmented generation for global story coherence and a dependency-aware visual consistency mechanism that tracks character and environmental states across temporal boundaries, while VLM-guided agents continuously monitor and refine both narrative coherence and visual fidelity. The framework enables coordinated agent collaboration to generate extended narrative content. This maintains both storytelling integrity and visual coherence across multi-scene timelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。