arXiv:2512.22536cs.CVcs.AI2025-12被引 5

CoAgent通过协作规划与验证,让长视频生成更连贯一致。

CoAgent: Collaborative Planning and Consistency Agent for Coherent Video Generation

  • 分步规划+全局记忆,确保角色和场景跨镜头一致
  • 生成中实时检测不一致并重生成,提升整体连贯性
  • 适合需要长视频叙事质量的生成任务

保持叙事连贯性和视觉一致性仍是开放域视频生成的核心挑战。现有文本到视频模型常独立处理每一镜头,导致身份漂移、场景不一致和时间结构不稳定。我们提出CoAgent,一种协同闭环框架,将生成过程构造成计划-合成-验证的流水线。给定用户提示、风格参考和节奏约束,一个分镜规划器将输入分解为带显式实体、空间关系和时间线索的结构化镜头计划。全局上下文管理器维护实体级记忆,以保持跨镜头的外观与身份一致性。每个镜头由合成模块在视觉一致性控制器引导下生成,同时验证代理使用视觉-语言推理评估中间结果,并在检测到不一致时触发选择性重生成。最后,节奏感知编辑器优化时间节奏与转场,匹配期望的叙事流。大量实验表明,CoAgent显著提升了长视频生成中的连贯性、视觉一致性和叙事质量。

原文摘要 · Abstract (English)

Maintaining narrative coherence and visual consistency remains a central challenge in open-domain video generation. Existing text-to-video models often treat each shot independently, resulting in identity drift, scene inconsistency, and unstable temporal structure. We propose CoAgent, a collaborative and closed-loop framework for coherent video generation that formulates the process as a plan-synthesize-verify pipeline. Given a user prompt, style reference, and pacing constraints, a Storyboard Planner decomposes the input into structured shot-level plans with explicit entities, spatial relations, and temporal cues. A Global Context Manager maintains entity-level memory to preserve appearance and identity consistency across shots. Each shot is then generated by a Synthesis Module under the guidance of a Visual Consistency Controller, while a Verifier Agent evaluates intermediate results using vision-language reasoning and triggers selective regeneration when inconsistencies are detected. Finally, a pacing-aware editor refines temporal rhythm and transitions to match the desired narrative flow. Extensive experiments demonstrate that CoAgent significantly improves coherence, visual consistency, and narrative quality in long-form video generation.

视频生成一致性协同规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。