融合视觉与触觉信息,提升机器人在视觉受限下的操作成功率
GelFusion: Enhancing Robotic Manipulation under Visual Constraints via Visuotactile Fusion
- 用双通道特征表示同时捕捉纹理几何与动态交互信息
- 在三种高接触任务中成功率达87.3%,优于基线模型
- 适合需要精细触觉反馈的机器人操控场景
视觉-触觉感知能提供丰富的接触信息,有助于缓解模仿学习在视觉受限条件(如视觉模糊或遮挡)下的性能瓶颈。然而,有效融合视觉与触觉模态仍具挑战。本文提出GelFusion框架,通过高分辨率GelSight传感器整合触觉反馈以增强策略学习。该框架采用以视觉为主导的交叉注意力融合机制,核心为双通道触觉特征表示,同步利用纹理-几何与动态交互特征。我们在三个高接触任务(表面擦拭、插销、易碎物体抓取放置)上评估了GelFusion,结果显示其在成功率上显著优于基线方法,验证了结构的有效性。
原文摘要 · Abstract (English)
Visuotactile sensing offers rich contact information that can help mitigate performance bottlenecks in imitation learning, particularly under vision-limited conditions, such as ambiguous visual cues or occlusions. Effectively fusing visual and visuotactile modalities, however, presents ongoing challenges. We introduce GelFusion, a framework designed to enhance policies by integrating visuotactile feedback, specifically from high-resolution GelSight sensors. GelFusion using a vision-dominated cross-attention fusion mechanism incorporates visuotactile information into policy learning. To better provide rich contact information, the framework's core component is our dual-channel visuotactile feature representation, simultaneously leveraging both texture-geometric and dynamic interaction features. We evaluated GelFusion on three contact-rich tasks: surface wiping, peg insertion, and fragile object pick-and-place. Outperforming baselines, GelFusion shows the value of its structure in improving the success rate of policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。