arXiv:2606.26410cs.CV2026-06

从单目视频学习隐式3D物理,让模型自动推演物体运动规律。

Neural Voxel Dynamics: Learning Implicit 3D Physics via Volumetric Feature Advection

论文配图:Neural Voxel Dynamics: Learning Implicit 3D Physics via Volumetric Feature Advection
图 1 · 摘自论文原文
  • 将视频特征升维到3D体素空间,用流场方式模拟物理演化。
  • 无需物理引擎内参,仅靠视频信号实现长期结构稳定与物理合理。
  • 适合研究通用动态世界模型、3D生成与物理推理的学者。

我们提出一种自监督框架,直接从视频衍生的监督信号中学习隐式3D物理动态。现有生成视频模型虽视觉保真度高,但缺乏3D几何基础,常导致物理不一致和物体恒存性失败。为此,我们将预测瓶颈从2D图像空间转移到经单目深度先验支撑的‘提升’3D体素潜空间。方法将视频联合嵌入预测架构(V-JEPA)的语义特征反投影至体素网格,通过体素特征输运学习动作条件下的状态转移算子,将物理建模为时空状态输运问题,即学习隐式3D物理。与依赖显式经典模拟器的先进混合模型不同,本架构在高维V-JEPA特征中隐式追踪物质状态,支持单一统一管道下异质现象(如刚体在流体中的运动)的涌现模拟。训练仅需端到端视频信号与动作条件,无需物理引擎内部状态、标签或代理模型,在多个基准(CLEVERER, PhysInOne, PhysGaia)上表现出良好的长期结构稳定性与物理合理性。我们认为该工作为仅通过被动观察单目视频构建通用动态世界模型开辟了可扩展路径。

原文摘要 · Abstract (English)

We present a self-supervised framework for learning implicit 3D physical dynamics directly from video-derived supervisory signals. While current generative video models achieve high visual fidelity, they lack a 3D geometric foundation, often resulting in physical inconsistencies and a failure to maintain object permanence. We address this by shifting the predictive bottleneck from 2D image space to a `lifted' 3D Volumetric Latent Space. Our method unprojects semantic features from a Video Joint-Embedding Predictive Architecture (V-JEPA) into a voxelized grid, grounded by monocular depth priors. This lifting enables a Volumetric Feature Advection to learn an action-conditioned transition operator that treats physics as a spatio-temporal state advection problem, i.e., learn implicit 3D physics. Unlike state-of-the-art hybrid models that rely on explicit classical simulators for training and/or inference, our architecture tracks material states implicitly within high-dimensional V-JEPA features. This allows for the emergent simulation of heterogeneous phenomena (e.g., rigid body motion in fluid flow) within a single, unified pipeline. Supervised solely via end-to-end video-derived signal plus action conditions, without access to physics engine internal states, labels, or surrogate models, our model demonstrates good long-term structural stability and physical plausibility on multiple benchmarks (CLEVERER, PhysInOne, PhysGaia). We believe that this work opens a scalable pathway toward general-purpose dynamic world models that internalize the 3D invariants of the physical world solely through passive observation of monocular videos.

3D物理视频生成隐式建模体素

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。