用视觉触觉融合构建可主动探索的3D物体表示,无需预训练。
Gaussian Process-Based Active Exploration Strategies in Vision and Touch
- 基于高斯过程融合视觉与触觉,构建带不确定性的距离场
- 通过增量观测迭代优化几何,实现对复杂形状的精准重建
- 适合在无先验的现实环境中进行自主感知与探索的机器人
机器人因缺乏先验知识,在非结构化环境中难以理解物体的形状、材质和语义属性。而人类则通过交互式多传感器探索来学习。本文提出将视觉与触觉观测融合为统一的高斯过程距离场(GPDF)表示,用于主动感知物体属性。虽主要关注几何,但该方法也展现出建模表面属性的潜力。GPDF利用点云、解析梯度与海森矩阵及表面不确定性估计编码有符号距离,这些是常见神经网络形状表示所缺乏的属性。通过点云构建距离函数,GPDF无需大规模预训练即可通过聚合观测逐步更新。从初始视觉形状估计出发,框架结合可微渲染的密集视觉测量与高不确定性区域的触觉测量,迭代优化几何结构。通过量化多模态不确定性,规划探索动作以最大化信息增益,恢复精确3D结构。真实机器人实验使用固定在桌面上的Franka Research 3机械臂,配备自定义DIGIT触觉传感器和Intel Realsense D435 RGBD相机。实验对象为静止放置于桌面的物体。为提升可扩展性,研究了诱导点法等高斯过程近似方法。该概率多模态融合方法实现了对复杂物体几何的主动探索与建图,未来或可扩展至非几何属性。
原文摘要 · Abstract (English)
Robots struggle to understand object properties like shape, material, and semantics due to limited prior knowledge, hindering manipulation in unstructured environments. In contrast, humans learn these properties through interactive multi-sensor exploration. This work proposes fusing visual and tactile observations into a unified Gaussian Process Distance Field (GPDF) representation for active perception of object properties. While primarily focusing on geometry, this approach also demonstrates potential for modeling surface properties beyond geometry. The GPDF encodes signed distance using point cloud, analytic gradient and Hessian, and surface uncertainty estimates, which are attributes that common neural network shape representation lack. By utilizing a point cloud to construct a distance function, GPDF does not need extensive pretraining on large datasets and can incorporate observations by aggregation. Starting with an initial visual shape estimate, the framework iteratively refines the geometry by integrating dense vision measurements using differentiable rendering and tactile measurements at uncertain surface regions. By quantifying multi-sensor uncertainties, it plans exploratory motions to maximize information gain for recovering precise 3D structures. For the real-world robot experiment, we utilize the Franka Research 3 robot manipulator, which is fixed on a table and has a customized DIGIT tactile sensor and an Intel Realsense D435 RGBD camera mounted on the end-effector. In these experiments, the robot explores the shape and properties of objects assumed to be static and placed on the table. To improve scalability, we investigate approximation methods like inducing point method for Gaussian Processes. This probabilistic multi-modal fusion enables active exploration and mapping of complex object geometries, extending potentially beyond geometry.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。