arXiv:2512.13604cs.CV2025-12被引 8

LongVie 2实现五分钟可控长视频生成,兼顾画面质量与时间连贯性。

LongVie 2: Multimodal Controllable Ultra-Long Video World Model

  • 分三阶段训练:多模态引导、降质感知输入、历史上下文对齐
  • 在100段高分辨率视频上达到最强长程可控性与视觉保真度
  • 适合研究长视频生成、世界模型构建的学者与开发者

基于预训练视频生成系统构建视频世界模型是迈向通用时空智能的重要却具挑战性的步骤。世界模型需具备可控性、长期视觉质量与时间一致性三大特性。为此,我们采用渐进式方法——先增强可控性,再扩展至长期高质量生成。提出LongVie 2,一个端到端自回归框架,分三阶段训练:(1) 多模态引导,融合密集与稀疏控制信号,提供隐式世界级监督,提升可控性;(2) 输入帧的退化感知训练,弥合训练与长期推理之间的差距,维持高质量视觉表现;(3) 历史上下文引导,对齐相邻片段间的上下文信息,确保时间一致性。我们进一步提出LongVGenBench,包含100个高分辨率一分钟视频的综合基准,覆盖多样真实与合成环境。大量实验表明,LongVie 2在长程可控性、时间连贯性与视觉保真度方面均达当前最优,支持长达五分钟的连续视频生成,标志着统一视频世界建模的重要进展。

原文摘要 · Abstract (English)

Building video world models upon pretrained video generation systems represents an important yet challenging step toward general spatiotemporal intelligence. A world model should possess three essential properties: controllability, long-term visual quality, and temporal consistency. To this end, we take a progressive approach-first enhancing controllability and then extending toward long-term, high-quality generation. We present LongVie 2, an end-to-end autoregressive framework trained in three stages: (1) Multi-modal guidance, which integrates dense and sparse control signals to provide implicit world-level supervision and improve controllability; (2) Degradation-aware training on the input frame, bridging the gap between training and long-term inference to maintain high visual quality; and (3) History-context guidance, which aligns contextual information across adjacent clips to ensure temporal consistency. We further introduce LongVGenBench, a comprehensive benchmark comprising 100 high-resolution one-minute videos covering diverse real-world and synthetic environments. Extensive experiments demonstrate that LongVie 2 achieves state-of-the-art performance in long-range controllability, temporal coherence, and visual fidelity, and supports continuous video generation lasting up to five minutes, marking a significant step toward unified video world modeling.

视频生成世界模型长视频可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。