arXiv:2410.04680cs.ROcs.CV2024-10ICRA被引 16

用深度不确定度指导机器人视觉与触觉的最优感知选择

Next Best Sense: Guiding Vision and Touch with FisherRF for 3D Gaussian Splatting

  • 结合语义对齐与损失优化,提升少视角3D高斯点云重建质量
  • 基于费舍尔信息扩展新方法,在真实机器人上实现在线主动感知
  • 适合做机器人三维感知、主动学习与多模态融合的研究者

我们提出一种基于3D高斯点云(3DGS)的机器人主动下一最佳视角与触觉姿态选择框架。3DGS作为新兴的显式三维场景表示方法,兼具照片级真实感与几何准确性,但在实际机器人应用中受限于视图数量,随机采样易导致视图重叠冗余。为此,我们构建端到端在线训练与主动选择流程,显著提升少视图条件下的3DGS性能。首先,通过引入分割任意模型2(SAM2)进行语义深度对齐,并辅以皮尔逊深度与表面法向损失,改善真实场景的颜色与深度重建效果;其次,将费舍尔信息法(FisherRF)扩展至基于深度不确定性选择视角与触觉姿态,实现在真实机器人系统上的实时3DGS训练中在线选择。我们在复杂机器人场景中验证了该方法在定性和定量上的显著改进。

原文摘要 · Abstract (English)

We propose a framework for active next best view and touch selection for robotic manipulators using 3D Gaussian Splatting (3DGS). 3DGS is emerging as a useful explicit 3D scene representation for robotics, as it has the ability to represent scenes in a both photorealistic and geometrically accurate manner. However, in real-world, online robotic scenes where the number of views is limited given efficiency requirements, random view selection for 3DGS becomes impractical as views are often overlapping and redundant. We address this issue by proposing an end-to-end online training and active view selection pipeline, which enhances the performance of 3DGS in few-view robotics settings. We first elevate the performance of few-shot 3DGS with a novel semantic depth alignment method using Segment Anything Model 2 (SAM2) that we supplement with Pearson depth and surface normal loss to improve color and depth reconstruction of real-world scenes. We then extend FisherRF, a next-best-view selection method for 3DGS, to select views and touch poses based on depth uncertainty. We perform online view selection on a real robot system during live 3DGS training. We motivate our improvements to few-shot GS scenes, and extend depth-based FisherRF to them, where we demonstrate both qualitative and quantitative improvements on challenging robot scenes. For more information, please see our project page at https://arm.stanford.edu/next-best-sense.

3D重建主动感知机器人多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。