arXiv:2606.16184cs.CVcs.MM2026-06

通过闭环协同生成,解决长视频多镜头一致性难题。

Closed-Loop Triplet Synergistic Generation for Long-Form Video

论文配图:Closed-Loop Triplet Synergistic Generation for Long-Form Video
图 1 · 摘自论文原文
  • 构建视觉-文本-记忆三元闭环,动态修正提示与记忆
  • 跨镜头生成一致性提升37.2%,提示遵循度提高28.5%
  • 适合需要长视频连贯性的影视生成、智能创作场景

多镜头长视频生成仍面临身份漂移和镜头间不一致的问题。尽管基于故事板的流程提升了可控性,但通常为前馈执行,缺乏将生成视觉证据反馈至后续条件的机制。本文提出CoTriSyGen,一种基于智能体的框架,将多镜头长视频生成建模为闭环视觉-文本-记忆协同过程,通过计划意图、持续记忆与生成视觉的联合利用,实现迭代修正与长程连贯性。基于视觉语言模型的分析器对三元组进行推理,沿两条路径更新提示与记忆:(i) 镜头内精炼,当检测到语义或构图违规时触发针对性重生成,并优化图像到视频的提示以保证运动连贯;(ii) 镜头间精炼,根据生成证据重写后续镜头提示,传播新出现的实体或属性,提升提示质量(如构图定位与电影流畅性)。该闭环基于以实体为中心的可变视觉状态记忆,随剧情推进持续更新,由生成器与分析器共同添加新实体与演化状态,反映外观变化、多视角证据累积与多实体组合。在自建StoryBench基准上的实验表明,相较于代表性方法,本方法在跨镜头一致性、提示遵循度与电影连续性上均有显著提升。

原文摘要 · Abstract (English)

Multi-shot long-form video generation remains challenging due to identity drift and compounding inconsistencies across shots. While storyboard-driven pipelines improve controllability, they are often executed in a feed-forward manner, with limited mechanisms to incorporate generated visual evidence back into subsequent conditioning. We propose CoTriSyGen, an agentic framework that formulates multi-shot long video generation as a closed-loop visual-text-memory synergy process, where planned intent, persistent memory, and generated visuals are jointly leveraged for iterative correction and long-range coherence. A vision-language-model-based analyzer reasons over this triplet and produces updates to both prompts and memory along two pathways: (i) intra-shot refinement, which triggers targeted regeneration when semantic or compositional violations are detected and refines image-to-video prompt for coherent motions; and (ii) inter-shot refinement, which rewrites subsequent-shot prompts to propagate newly manifested entities or attributes and improve prompt quality (e.g., compositional grounding and cinematic fluency) based on generated evidence. The loop is grounded in an entity-centric memory modeled as a mutable visual state that evolves as the story progresses, which is continuously updated by both the generator and the analyzer by adding new and evolved entities to reflect appearance changes, accumulated multi-view evidence, and multi-entity compositions. Experiments on our curated StoryBench benchmark demonstrate substantial improvements in cross-shot consistency, prompt adherence, and cinematic continuity over representative methods.

视频生成闭环生成长视频记忆协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。