arXiv:2602.00148cs.CVcs.AI2026-02被引 5

用神经高斯力场实现快速物理驱动的4D视频生成

Learning Physics-Grounded 4D Dynamics with Neural Gaussian Force Fields

  • 将3D高斯点云与物理引擎结合,端到端生成动态4D视频
  • 速度比之前高斯模拟器快100倍,支持复杂物体交互
  • 适合需要物理真实性的视频预测、数字孪生等场景

从原始视觉数据中预测物理动态仍是人工智能的重大挑战。尽管近期视频生成模型在视觉质量上取得进展,但因缺乏对物理规律的建模,仍难以持续生成符合物理规律的视频。现有结合3D高斯溅射与物理引擎的方法虽能生成物理合理视频,但重建和仿真计算成本高,且在复杂真实场景中鲁棒性不足。为此,我们提出神经高斯力场(NGFF),一种将3D高斯感知与物理动力学建模融合的端到端神经框架,可从多视角RGB输入生成可交互的物理真实4D视频,速度比先前高斯模拟器快两个数量级。为支持训练,我们还构建了GSCollision数据集,包含超过64万帧渲染的物理视频(约4TB),涵盖多种材料、多物体交互及复杂场景。在合成与真实3D场景上的评估表明,NGFF在物理推理方面具有强泛化能力与鲁棒性,推动视频预测向物理基础世界模型发展。

原文摘要 · Abstract (English)

Predicting physical dynamics from raw visual data remains a major challenge in AI. While recent video generation models have achieved impressive visual quality, they still cannot consistently generate physically plausible videos due to a lack of modeling of physical laws. Recent approaches combining 3D Gaussian splatting and physics engines can produce physically plausible videos, but are hindered by high computational costs in both reconstruction and simulation, and often lack robustness in complex real-world scenarios. To address these issues, we introduce Neural Gaussian Force Field (NGFF), an end-to-end neural framework that integrates 3D Gaussian perception with physics-based dynamic modeling to generate interactive, physically realistic 4D videos from multi-view RGB inputs, achieving two orders of magnitude faster than prior Gaussian simulators. To support training, we also present GSCollision, a 4D Gaussian dataset featuring diverse materials, multi-object interactions, and complex scenes, totaling over 640k rendered physical videos (~4 TB). Evaluations on synthetic and real 3D scenarios show NGFF's strong generalization and robustness in physical reasoning, advancing video prediction towards physics-grounded world models.

4D视频生成物理建模高斯溅射动态预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。