用视觉思维链提升视频生成的逻辑连贯性
VChain: Chain-of-Visual-Thought for Reasoning in Video Generation
- 引入多模态模型生成关键帧,引导视频生成过程
- 在关键帧处稀疏调整状态,显著提升复杂场景生成质量
- 无需微调,适合需要高逻辑一致性的视频生成任务
近期视频生成模型虽能产出流畅自然的视频片段,但在合成具有连贯因果关系的复杂动态时仍表现不佳。准确建模随时间演变的视觉状态与结果仍是核心挑战。相比之下,大语言与多模态模型(如 GPT-4o)具备强大的视觉状态推理与未来预测能力。为此,我们提出 VChain——一种基于推理时视觉思维链的新框架,将多模态模型的视觉推理信号注入视频生成流程。VChain 设计了专用管道,利用大范围多模态模型生成稀疏的关键帧作为快照,并仅在这些关键点对预训练视频生成器进行稀疏的推理时状态适应。该方法无需参数微调,开销小,避免密集监督。在复杂多步场景下的大量实验表明,VChain 显著提升了生成视频的质量。
原文摘要 · Abstract (English)
Recent video generation models can produce smooth and visually appealing clips, but they often struggle to synthesize complex dynamics with a coherent chain of consequences. Accurately modeling visual outcomes and state transitions over time remains a core challenge. In contrast, large language and multimodal models (e.g., GPT-4o) exhibit strong visual state reasoning and future prediction capabilities. To bridge these strengths, we introduce VChain, a novel inference-time chain-of-visual-thought framework that injects visual reasoning signals from multimodal models into video generation. Specifically, VChain contains a dedicated pipeline that leverages large multimodal models to generate a sparse set of critical keyframes as snapshots, which are then used to guide the sparse inference-time visual-state adaptation of a pre-trained video generator only at these key moments. Our approach is tuning-efficient, introduces minimal overhead and avoids dense supervision. Extensive experiments on complex, multi-step scenarios show that VChain significantly enhances the quality of generated videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。