arXiv:2604.09877cs.CVcs.AI2026-04

用手机拍摄生成可交互的语义化动态4D场景模型

Genie 4D: Semantic-Prior-Guided 4D Dynamic Scene Reconstruction

论文配图:Genie 4D: Semantic-Prior-Guided 4D Dynamic Scene Reconstruction
图 1 · 摘自论文原文
  • 结合视觉惯性与冻结的DINOv3语义特征,实现动态追踪中的结构先验约束
  • 在Point Odyssey和TUM-Dynamics上提升3D跟踪精度与重建完整性,保持线性时间复杂度
  • 支持用户动作控制,可在消费级显卡上实时运行,适合移动设备和多平台应用

在计算机视觉与机器人感知交汇处,动态场景的4D重建将底层几何感知与高层语义理解相连接。我们提出Genie 4D,一个将手持手机拍摄转化为语义锚定、动作可控的4D世界模型的框架。Genie 4D采用实时视觉-惯性高斯点云前端获取度量几何,搭配由冻结的DINOv3特征作为结构先验的前馈式4D主干网络。语义先验抑制了动态追踪中的身份漂移,短时条件扩散重构器恢复了回归主干所平滑掉的高频表面细节。最终,轻量级潜在动作头将重建的4D状态暴露给以JEPA风格下一项嵌入目标训练的类Genie世界模型,使场景可在用户操作下向前推演。在Point Odyssey和TUM-Dynamics基准上,Genie 4D在保持前馈基线线性时间复杂度O(T)的同时,提升了3D跟踪准确率(APD)和重建完整度,并可在单张消费级显卡(RTX 5090)上从iPhone、Mac、Windows和Linux客户端实现交互式运行。Genie 4D为物理一致的世界模型提供了一条实用的、基于语义先验的路径。

原文摘要 · Abstract (English)

At the intersection of computer vision and robotic perception, 4D reconstruction of dynamic scenes connects low-level geometric sensing with high-level semantic understanding. We present Genie 4D, a framework that turns hand-held phone capture into a semantically grounded, action-controllable 4D world model. Genie 4D couples a real-time visual-inertial Gaussian splatting front-end for metric geometry with a feed-forward 4D backbone regularized by frozen DINOv3 features acting as structural priors. The semantic priors suppress identity drift during dynamic tracking, while a short conditional diffusion refiner recovers high-frequency surface detail that regression backbones smooth away. Finally, a lightweight latent-action head exposes the reconstructed 4D state to a Genie-style world model trained with a JEPA-style next-embedding objective, so that the scene can be rolled forward under user actions. On the Point Odyssey and TUM-Dynamics benchmarks, Genie 4D retains the linear time complexity O(T) of feed-forward baselines while improving 3D tracking accuracy (APD) and reconstruction completeness, and it runs interactively on a single consumer GPU (RTX 5090) from iPhone, Mac, Windows, and Linux capture clients. Genie 4D offers a practical, semantic-prior-guided path toward physically grounded world models.

4D重建语义先验动态场景可交互建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。