arXiv:2605.25909cs.CV2026-05

用刚体约束提升动态场景重建效率,支持文本查询物体

R5DGS: Semantic-Aware 4D Gaussian Splatting with Rigid Body Constraints for Efficient Dynamic Scene Reconstruction

论文配图:R5DGS: Semantic-Aware 4D Gaussian Splatting with Rigid Body Constraints for Efficient Dynamic Scene Reconstruction
图 1 · 摘自论文原文
  • 引入身份编码向量关联高斯点与物体,实现语义感知
  • 通过刚体约束使预测速度提升11倍,轨迹仍合理
  • 支持任意视角和时间点的文本驱动物体渲染

从多视角视频中重建并预测动态三维场景是机器人、AR/VR和数字孪生的基础任务。现有物理引导的高斯点渲染方法虽能良好外推未来帧,但缺乏语义感知且计算开销大。我们提出R5DGS框架,在物理驱动的4D高斯表示中加入紧凑的身份编码向量,实现高斯点到物体的精确关联。通过离线构建基于CLIP的物体查找表,支持开放词汇文本提示,可在任意时间戳和视角下检索并渲染特定物体的高斯点。此外,我们提出刚体推理约束,仅对物体质心进行物理动力学预测,并通过相对变换传播运动至相关高斯点。该优化在外推过程中实现11 FPS的速度提升,同时保持轨迹合理性。

原文摘要 · Abstract (English)

Reconstructing and predicting dynamic 3D scenes from multi-view videos is a foundational task for robotics, AR/VR, and digital twins. Recent physics-informed Gaussian Splatting methods achieve impressive future frame extrapolation but lack semantic awareness and suffer from large computational overhead. We introduce $\textbf{R5DGS}$, a framework that augments a physics-driven 4D Gaussian representation with compact Identity Encoding vectors, enabling precise Gaussian-to-object association. By constructing an offline CLIP-based object lookup table, we support open-vocabulary text prompting to retrieve and render object-specific Gaussians across arbitrary timestamps and viewpoints. Furthermore, we propose a rigid-body inference constraint that predicts and integrates physical dynamics exclusively for object centroids, propagating motion to associated Gaussians via relative transformations. This optimization yields a 11 FPS speedup during extrapolation without compromising trajectories plausibility.

3D重建高斯溅射动态场景语义感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。