arXiv:2607.26452cs.AIcs.CV2026-07

构建大规模世界状态数据集,支持智能体的因果推理与可控生成。

CG-World: A Large-Scale World-State Dataset and Protocol for World Models

论文配图:CG-World: A Large-Scale World-State Dataset and Protocol for World Models
图 1 · 摘自论文原文
  • 从工业级渲染管线提取多模态中间状态,结构化组织时空数据。
  • 包含约85万段1-5秒的对齐片段,支持干预与反事实分析。
  • 适合研究世界模型、物理人工智能和具身智能的开发者使用。

世界模型需学习状态、动作、事件与观测的联合动态,但现有视频、机器人及仿真数据集通常仅覆盖部分结构。我们提出CG-World,一个源自工业计算机图形制作流程的大规模世界状态数据集与协议。该数据集显式记录中间状态,包括多模态语义、空间结构、骨骼与控制器状态、运动曲线、相机与光照参数、物理缓存、接触事件及多通道渲染结果。CG-World v1包含约85万条1-5秒的时序对齐片段,分离潜变量状态、观测、关系、事件与分支元数据,并整合为统一的时空样本。为支持干预学习与反事实推理,定义了涵盖真实轨迹、观测干预、动作干预、机制干预与严格反事实分支的分支谱系,明确标注干预目标、不变量与替代结果。我们在几何条件视频生成、动作预测与闭环视觉-语言-动作策略迁移任务上评估该数据集,结果表明其能为受控生成、动作建模与具身策略迁移提供可复用的结构化监督。未来计划通过持续收集与社区协作扩展数据集,共建世界模型、物理人工智能与具身智能的共享数据基础设施。

原文摘要 · Abstract (English)

World models must learn the joint dynamics of states, actions, events, and observations, yet existing video, robotics, and simulation datasets usually capture only part of this structure. We introduce CG-World, a large-scale world-state dataset and protocol derived from industrial computer graphics production pipelines. CG-World explicitly records intermediate states, including multimodal semantics, spatial structure, skeletal and controller states, motion curves, camera and lighting parameters, physics caches, contact events, and multi-pass renderings. CG-World v1 contains approximately 850,000 temporally aligned segments of 1-5 seconds. It separates latent states, observations, relations, events, and branch metadata, and organizes them into unified spatiotemporal samples. To support intervention learning and counterfactual reasoning, CG-World defines a branch lineage covering factual trajectories, observation interventions, action interventions, mechanism interventions, and strict counterfactual branches, with intervention targets, invariants, and alternative outcomes explicitly recorded. We evaluate the dataset on geometry-conditioned video generation, action prediction, and closed-loop vision-language-action policy transfer. Results show that CG-World provides reusable structured supervision for controlled generation, action modeling, and embodied policy transfer. We plan to expand CG-World through continued data collection and community collaboration toward a shared data infrastructure for world models, Physical AI, and embodied intelligence.

世界模型数据集物理AI具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。