arXiv:2512.03000cs.CV2025-12NeurIPS被引 11

构建可感知物理尺度的4D动态世界模型,提升智能体对真实场景的理解能力。

DynamicVerse: A Physically-Aware Multimodal Framework for 4D World Modeling

  • 融合视觉、几何与多模态模型,从单目视频中重建物理尺度的动态结构
  • 生成超10万段视频、80万标注掩码、超1000万帧的4D数据集
  • 适用于具身智能、自动驾驶等需理解真实动态环境的任务

理解具有演变3D结构、真实运动及文本描述的动态物理世界,对人机交互至关重要,使智能体具备类人感知与行动能力。然而现有数据集多源于有限模拟器或传统SfM方法,标注规模小且描述性字幕匮乏,限制了基础模型从互联网单目视频中准确解析真实动态的能力。为此,我们提出DynamicVerse——一个面向真实世界视频的物理尺度多模态4D建模框架。通过大型视觉、几何与多模态模型,该框架可解析度量尺度的静态几何、真实动态运动、实例级掩码及整体描述性文字。结合窗口式束调整与全局优化,将长时真实视频序列转换为完整的4D多模态格式。DynamicVerse构建了包含10万+视频、80万+标注掩码和1000万+帧的大规模数据集。在视频深度估计、相机位姿估计与内参估计三个基准任务上的实验表明,该4D建模在捕捉物理尺度测量上优于现有方法,具有更高的全局精度。

原文摘要 · Abstract (English)

Understanding the dynamic physical world, characterized by its evolving 3D structure, real-world motion, and semantic content with textual descriptions, is crucial for human-agent interaction and enables embodied agents to perceive and act within real environments with human-like capabilities. However, existing datasets are often derived from limited simulators or utilize traditional Structurefrom-Motion for up-to-scale annotation and offer limited descriptive captioning, which restricts the capacity of foundation models to accurately interpret real-world dynamics from monocular videos, commonly sourced from the internet. To bridge these gaps, we introduce DynamicVerse, a physical-scale, multimodal 4D world modeling framework for dynamic real-world video. We employ large vision, geometric, and multimodal models to interpret metric-scale static geometry, real-world dynamic motion, instance-level masks, and holistic descriptive captions. By integrating window-based Bundle Adjustment with global optimization, our method converts long real-world video sequences into a comprehensive 4D multimodal format. DynamicVerse delivers a large-scale dataset consisting of 100K+ videos with 800K+ annotated masks and 10M+ frames from internet videos. Experimental evaluations on three benchmark tasks, namely video depth estimation, camera pose estimation, and camera intrinsics estimation, demonstrate that our 4D modeling achieves superior performance in capturing physical-scale measurements with greater global accuracy than existing methods.

4D建模多模态动态世界视觉几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。