arXiv:2410.15730cs.RO2024-10被引 2

用2D高斯表示动态场景,实现语义、运动与几何的统一建模。

MSGField: A Unified Scene Representation Integrating Motion, Semantics, and Geometry for Robotic Manipulation

  • 用有限运动基分解物体运动,紧凑编码动态信息。
  • 双视角图像监督下实时优化复杂非刚性运动,静态成功率79.2%。
  • 适合小物体与柔性物体操控,适用于真实机器人任务。

将精确几何与丰富语义结合已被证明在语言引导的机器人操作中非常有效。现有动态场景方法或无法实时更新,或依赖额外深度传感器进行简单编辑,限制了其在真实场景中的应用。本文提出MSGField,利用一组2D高斯实现高质量重建,并通过属性编码语义与运动信息。特别地,通过将每个物体的运动分解为一组有限的运动基,实现运动场的紧凑表示。借助高斯点云的可微实时渲染,仅需两个相机视角的图像监督,即可快速优化物体运动,即使面对复杂非刚性运动。此外,设计了基于物体先验的管道,高效获取明确语义。在包含柔性及极小物体的挑战性数据集上,静态环境下语言引导操作成功率达79.2%,动态环境达63.3%;针对特定物体抓取,成功率90%,媲美基于点云的方法。代码与数据集将公开于:https://shengyu724.github.io/MSGField.github.io。

原文摘要 · Abstract (English)

Combining accurate geometry with rich semantics has been proven to be highly effective for language-guided robotic manipulation. Existing methods for dynamic scenes either fail to update in real-time or rely on additional depth sensors for simple scene editing, limiting their applicability in real-world. In this paper, we introduce MSGField, a representation that uses a collection of 2D Gaussians for high-quality reconstruction, further enhanced with attributes to encode semantic and motion information. Specially, we represent the motion field compactly by decomposing each primitive's motion into a combination of a limited set of motion bases. Leveraging the differentiable real-time rendering of Gaussian splatting, we can quickly optimize object motion, even for complex non-rigid motions, with image supervision from only two camera views. Additionally, we designed a pipeline that utilizes object priors to efficiently obtain well-defined semantics. In our challenging dataset, which includes flexible and extremely small objects, our method achieve a success rate of 79.2% in static and 63.3% in dynamic environments for language-guided manipulation. For specified object grasping, we achieve a success rate of 90%, on par with point cloud-based methods. Code and dataset will be released at:https://shengyu724.github.io/MSGField.github.io.

机器人操作动态场景高斯表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。