DreamWorld让视频生成具备统一的世界理解能力,提升时空一致性。
DreamWorld: Unified World Modeling in Video Generation
- 通过联合建模物理常识、3D结构与时间动态,构建统一世界模型
- 在VBench上比Wan2.1提升2.26分,显著改善视频连贯性
- 适合需要高质量长视频生成的研究者与开发者
尽管视频生成取得显著进展,现有模型仍局限于表面合理性,缺乏对世界的统一理解。以往方法通常仅引入单一形式的世界知识或依赖固定对齐策略,难以实现多维度(如物理常识、3D一致性、时间连续性)的协同建模。为此,我们提出DreamWorld,一种基于联合世界建模范式的统一框架,通过从基础模型中联合预测视频像素与特征,捕捉时序动态、空间几何与语义一致性。然而,直接优化这些异构目标易导致视觉不稳与时间闪烁。为此,我们设计一致约束渐进调节(CCA)以逐步调控世界级约束,并引入多源内引导机制,在推理阶段强化学习到的世界先验。大量实验表明,DreamWorld在世界一致性方面显著优于Wan2.1,在VBench上提升2.26分。代码将公开于Github。
原文摘要 · Abstract (English)
Despite impressive progress in video generation, existing models remain limited to surface-level plausibility, lacking a coherent and unified understanding of the world. Prior approaches typically incorporate only a single form of world-related knowledge or rely on rigid alignment strategies to introduce additional knowledge. However, aligning the single world knowledge is insufficient to constitute a world model that requires jointly modeling multiple heterogeneous dimensions (e.g., physical commonsense, 3D and temporal consistency). To address this limitation, we introduce \textbf{DreamWorld}, a unified framework that integrates complementary world knowledge into video generators via a \textbf{Joint World Modeling Paradigm}, jointly predicting video pixels and features from foundation models to capture temporal dynamics, spatial geometry, and semantic consistency. However, naively optimizing these heterogeneous objectives can lead to visual instability and temporal flickering. To mitigate this issue, we propose \textit{Consistent Constraint Annealing (CCA)} to progressively regulate world-level constraints during training, and \textit{Multi-Source Inner-Guidance} to enforce learned world priors at inference. Extensive evaluations show that DreamWorld improves world consistency, outperforming Wan2.1 by 2.26 points on VBench. Code will be made publicly available at \href{https://github.com/ABU121111/DreamWorld}{\textcolor{mypink}{\textbf{Github}}}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。