解决超长视频生成中不连贯与画质下降问题,支持多模态可控生成。
LongVie: Multimodal-Guided Controllable Ultra-Long Video Generation
- 统一噪声初始化与全局控制归一化,保障视频时序一致性。
- 融合稠密与稀疏多模态信号,生成1分钟以上高质量视频。
- 提出新基准LongVGenBench,验证长视频可控生成性能领先。
可控的超长视频生成是基础但极具挑战的任务。现有方法在短片段上有效,但在扩展时面临时间不一致和视觉退化问题。本文首次分析并识别出三个关键因素:独立噪声初始化、独立控制信号归一化以及单模态引导的局限性。为此,我们提出LongVie,一个端到端自回归框架,用于可控长视频生成。LongVie引入两项核心设计以确保时序一致性:1)统一噪声初始化策略,维持各片段间生成一致性;2)全局控制信号归一化,实现全视频控制空间对齐。为缓解视觉退化,LongVie采用3)多模态控制框架,整合深度图等稠密信号与关键点等稀疏信号,并结合4)退化感知训练策略,动态调节模态贡献以保持视觉质量。我们还构建了LongVGenBench,包含100个高分辨率视频,覆盖多样真实与合成环境,每段超过一分钟。大量实验表明,LongVie在长程可控性、一致性和质量上均达到当前最优水平。
原文摘要 · Abstract (English)
Controllable ultra-long video generation is a fundamental yet challenging task. Although existing methods are effective for short clips, they struggle to scale due to issues such as temporal inconsistency and visual degradation. In this paper, we initially investigate and identify three key factors: separate noise initialization, independent control signal normalization, and the limitations of single-modality guidance. To address these issues, we propose LongVie, an end-to-end autoregressive framework for controllable long video generation. LongVie introduces two core designs to ensure temporal consistency: 1) a unified noise initialization strategy that maintains consistent generation across clips, and 2) global control signal normalization that enforces alignment in the control space throughout the entire video. To mitigate visual degradation, LongVie employs 3) a multi-modal control framework that integrates both dense (e.g., depth maps) and sparse (e.g., keypoints) control signals, complemented by 4) a degradation-aware training strategy that adaptively balances modality contributions over time to preserve visual quality. We also introduce LongVGenBench, a comprehensive benchmark consisting of 100 high-resolution videos spanning diverse real-world and synthetic environments, each lasting over one minute. Extensive experiments show that LongVie achieves state-of-the-art performance in long-range controllability, consistency, and quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。