arXiv:2410.08282cs.ROcs.AI2024-10ICRA被引 8

融合常识与视觉触觉,实现稀疏视角下鲁棒3D重建

FusionSense: Bridging Common Sense, Vision, and Touch for Robust Sparse-View Reconstruction

  • 用3D高斯点云结合基础模型先验,构建全局形状
  • 在透明/反光/暗色物体上仍保持快速稳定重建
  • 适合需精准感知的机器人抓取与导航任务

人类能自然融合常识、视觉和触觉来理解环境。为此,我们提出FusionSense——一种新型3D重建框架,使机器人能够融合基础模型的先验知识与来自视觉和触觉传感器的稀疏观测数据。该框架解决三大挑战:(i) 如何高效获取周围场景和物体的鲁棒全局形状信息?(ii) 如何利用几何与常识先验策略性选择触觉接触点?(iii) 如何通过部分触觉信号提升整体物体表征?系统以3D高斯点云为核心表示,采用分层优化策略,包括全局结构构建、物体视觉外壳裁剪与局部几何约束。该方法在传统难处理物体(如透明、反光或暗色)上仍实现快速且稳定的感知,显著提升下游操作与导航能力。真实世界实验表明,该框架优于现有最先进稀疏视角重建方法。所有代码与数据已在项目网站开源。

原文摘要 · Abstract (English)

Humans effortlessly integrate common-sense knowledge with sensory input from vision and touch to understand their surroundings. Emulating this capability, we introduce FusionSense, a novel 3D reconstruction framework that enables robots to fuse priors from foundation models with highly sparse observations from vision and tactile sensors. FusionSense addresses three key challenges: (i) How can robots efficiently acquire robust global shape information about the surrounding scene and objects? (ii) How can robots strategically select touch points on the object using geometric and common-sense priors? (iii) How can partial observations such as tactile signals improve the overall representation of the object? Our framework employs 3D Gaussian Splatting as a core representation and incorporates a hierarchical optimization strategy involving global structure construction, object visual hull pruning and local geometric constraints. This advancement results in fast and robust perception in environments with traditionally challenging objects that are transparent, reflective, or dark, enabling more downstream manipulation or navigation tasks. Experiments on real-world data suggest that our framework outperforms previously state-of-the-art sparse-view methods. All code and data are open-sourced on the project website.

3D重建多模态感知机器人感知高斯点云

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。