提出世界模型三重一致性原则,为通用智能提供理论框架。
The Trinity of Consistency as a Defining Principle for General World Models
- 以模态、空间、时间三一致为核心构建世界模型
- 提出CoW-Bench基准,统一评估多帧生成与推理能力
- 适合研究通用智能与多模态系统架构的学者
构建能够学习、模拟和推理客观物理规律的世界模型,是实现通用人工智能的基础挑战。尽管以Sora为代表的视频生成模型展示了数据驱动扩展法则在逼近物理动态方面的潜力,而新兴的统一多模态模型(UMM)也提供了融合感知、语言与推理的有前景架构,但该领域仍缺乏定义通用世界模型必要属性的原则性理论框架。本文提出,世界模型必须基于三重一致性:模态一致性作为语义接口,空间一致性作为几何基础,时间一致性作为因果引擎。通过这一三元视角,我们系统回顾了多模态学习的发展历程,揭示出从松散耦合的专业模块向统一架构演进的趋势,推动内部世界模拟器的协同涌现。为补充该概念框架,我们引入CoW-Bench基准,聚焦多帧推理与生成场景,采用统一评估协议测试视频生成模型与UMMs。本工作确立了通向通用世界模型的原则路径,阐明了当前系统的局限性及未来发展的架构要求。
原文摘要 · Abstract (English)
The construction of World Models capable of learning, simulating, and reasoning about objective physical laws constitutes a foundational challenge in the pursuit of Artificial General Intelligence. Recent advancements represented by video generation models like Sora have demonstrated the potential of data-driven scaling laws to approximate physical dynamics, while the emerging Unified Multimodal Model (UMM) offers a promising architectural paradigm for integrating perception, language, and reasoning. Despite these advances, the field still lacks a principled theoretical framework that defines the essential properties requisite for a General World Model. In this paper, we propose that a World Model must be grounded in the Trinity of Consistency: Modal Consistency as the semantic interface, Spatial Consistency as the geometric basis, and Temporal Consistency as the causal engine. Through this tripartite lens, we systematically review the evolution of multimodal learning, revealing a trajectory from loosely coupled specialized modules toward unified architectures that enable the synergistic emergence of internal world simulators. To complement this conceptual framework, we introduce CoW-Bench, a benchmark centered on multi-frame reasoning and generation scenarios. CoW-Bench evaluates both video generation models and UMMs under a unified evaluation protocol. Our work establishes a principled pathway toward general world models, clarifying both the limitations of current systems and the architectural requirements for future progress.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。