用视觉触觉连续反馈提升机器人灵巧操作能力
FingerEye: Learning Dexterous Manipulation with Continuous Vision-Tactile Sensing
- 融合双目相机与柔性接触层,实现接触前后全程感知
- 真实与仿真环境下任务成功率提升超30个百分点
- 适合研究灵巧抓取、多模态感知的科研人员
灵巧机器人操作需要从接触前到接触后都保持有效的感知。我们提出FingerEye,一个通过交互全程的视觉-触觉反馈来增强机器人灵巧性的感知与学习框架。传感方面,FingerEye将双目RGB相机与柔性接触界面结合,接触前由指尖摄像头提供近距离视觉线索和隐式立体信息,用于精确逼近与物体定位;接触后,通过标记点追踪柔性环的形变,作为接触力矩的代理信号。学习方面,构建了真实与仿真并行的数据采集与评估基础设施,系统研究了多FingerEye传感器下的策略-接口设计,并提出FingerEye Policy,采用分组结构化的模态融合方法,减少模态偏差,更好利用分布式指尖反馈。在七个接触敏感任务中,FingerEye在仿真和真实世界中均使仅使用腕部控制的策略平均成功率提升超过30个百分点。
原文摘要 · Abstract (English)
Dexterous robotic manipulation requires perception that remains informative from pre-contact approach to contact initiation and post-contact control. We introduce FingerEye, a sensing and learning framework that strengthens robotic dexterity through continuous vision-tactile feedback throughout interaction. On the sensing side, FingerEye integrates binocular RGB cameras with a compliant contact interface to support perception both before and after contact. Before contact, the fingertip cameras provide close-range visual cues and implicit stereo for precise approach and object localization. After contact, marker-tracked deformation of the compliant ring provides a proxy for contact wrench sensing. On the learning side, we build real-and-sim infrastructure for data collection and evaluation, systematically study policy-interface designs for learning with multiple FingerEye sensors, and develop FingerEye Policy, which applies group-structured modality fusion to reduce modality shortcuts and better exploit distributed fingertip feedback. Across seven contact-sensitive task settings, FingerEye improves wrist-only policy by over 30 percentage points in mean success rate in both simulation and the real world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。