arXiv:2606.31388cs.CV2026-06中稿 · ECCV被引 3

单视频生成可模拟的4D物理场景,无需训练

One Video, One World: Turning Monocular Video into Physical 4D Scenes

论文配图:One Video, One World: Turning Monocular Video into Physical 4D Scenes
图 1 · 摘自论文原文
  • 用视觉语言模型识别物体并分类运动,直接通过顶点变形重建
  • 在两个合成基准上精度领先,速度比基线快10到100倍
  • 输出结果可直接用于物理仿真,适合做智能体环境建模

我们提出OVOW,首个无需训练即可从单视角视频重建实例级、可模拟4D网格场景的系统。现有4D重建虽渲染质量高,但输出(如隐式场、高斯原语或点云)缺乏物理仿真所需的封闭拓扑、实例分离和标准接口。OVOW采用四阶段流程:视觉语言模型发现、标注并分类所有物体;类别感知重建生成刚体物体的独立网格及可变形物体的拓扑一致网格序列;迭代渲染-匹配-优化恢复度量尺度与6-DoF位姿轨迹;物理驱动组装确保地面接触与物体间支撑关系。关键在于,所有运动(刚性与非刚性)均通过直接顶点变形建模,无需类别特定先验或骨骼绑定,生成可直接用于下游物理仿真与编辑的封闭网格场景。我们还建立了首个结构化视频转4D评估基准,包含几何正确性、实例分离与物理合理性指标,超越视觉保真度;同一管道还可作为可扩展引擎,用于未来4D世界模型与具身智能的视频-4D配对数据合成。在两个合成基准(静态与4D)上,OVOW在布局与几何精度上表现最优,光度与语义误差最低,且单目视频运行速度比基线快1至2个数量级,下游物理仿真验证了其物理稳定性。

原文摘要 · Abstract (English)

We introduce \textbf{OVOW}, the first training-free system that reconstructs \emph{instance-level, simulation-ready} 4D mesh scenes from a single monocular video. Recent 4D reconstruction achieves impressive rendering quality, but its outputs (\eg, implicit fields, Gaussian primitives, or point clouds) lack the watertight topology, instance separation, and standardized physical interfaces required by physics simulators and embodied AI. OVOW closes this gap with a four-stage pipeline: a vision-language model discovers, labels, and motion-classifies all instances; category-aware reconstruction yields per-instance meshes for rigid objects and topology-consistent mesh sequences for deformable ones; an iterative render-match-optimize procedure recovers metric scale and 6-DoF pose trajectories; and physics-grounded assembly enforces ground contact and inter-object support. Crucially, we model all motion, rigid and non-rigid, through direct vertex deformation without category-specific priors or skeleton rigging, producing watertight mesh scenes ready for downstream physics simulation and editing. We further establish the first benchmark for \emph{structured Video-to-4D} evaluation, with metrics for geometric correctness, instance separation, and physical plausibility beyond visual fidelity; the same pipeline doubles as a scalable engine for \emph{synthesizing} paired video-to-4D simulation data for future 4D world models and embodied AI. Across two synthetic benchmarks (static and 4D), OVOW attains the best overall layout and geometry accuracy and the lowest photometric and semantic error among all baselines, and on monocular video runs one to two orders of magnitude faster than the baselines, while downstream physics simulation confirms its physical stability.

4D重建物理仿真单视频重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。