arXiv:2604.09415cs.CVcs.AI2026-04被引 7

构建200万视频的物理数据集,助力AI理解真实世界运动规律。

PhysInOne: Visual Physics Learning and Reasoning in One Suite

  • 生成200万视频覆盖71种物理现象,含多物体复杂交互。
  • 在4类任务中验证效果,显著提升生成视频的物理合理性。
  • 适合研究物理推理、世界模型与具身智能的开发者使用。

我们提出PhysInOne,一个大规模合成数据集,解决人工智能系统在物理基础训练数据上的严重匮乏问题。不同于现有仅数百或数千样本的数据集,PhysInOne包含200万视频,覆盖153,810个动态3D场景,涵盖力学、光学、流体动力学和磁学中的71种基本物理现象。其场景具备复杂背景下的多物体交互,提供全面的真值标注,包括3D几何、语义、动态运动、物理属性及文本描述。我们在四个新兴应用中验证了PhysInOne的有效性:物理感知视频生成、长/短期未来帧预测、物理属性估计和运动迁移。实验表明,基于PhysInOne微调基础模型可显著提升生成结果的物理合理性,同时揭示了建模复杂物理动态和估计内在属性的关键缺陷。作为同类中规模最大的数据集,其规模远超以往工作,为生成、仿真与具身智能中的物理基础世界模型发展树立了新基准。

原文摘要 · Abstract (English)

We present PhysInOne, a large-scale synthetic dataset addressing the critical scarcity of physically-grounded training data for AI systems. Unlike existing datasets limited to merely hundreds or thousands of examples, PhysInOne provides 2 million videos across 153,810 dynamic 3D scenes, covering 71 basic physical phenomena in mechanics, optics, fluid dynamics, and magnetism. Distinct from previous works, our scenes feature multiobject interactions against complex backgrounds, with comprehensive ground-truth annotations including 3D geometry, semantics, dynamic motion, physical properties, and text descriptions. We demonstrate PhysInOne's efficacy across four emerging applications: physics-aware video generation, long-/short-term future frame prediction, physical property estimation, and motion transfer. Experiments show that fine-tuning foundation models on PhysInOne significantly enhances physical plausibility, while also exposing critical gaps in modeling complex physical dynamics and estimating intrinsic properties. As the largest dataset of its kind, orders of magnitude beyond prior works, PhysInOne establishes a new benchmark for advancing physics-grounded world models in generation, simulation, and embodied AI.

物理建模世界模型视频生成具身智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。