用触觉信号提升手部抓握时被遮挡物体的6D姿态追踪精度。
KineFuse: Kinematic-Aware Haptic Fusion for In-Hand Occluded-Object Pose Tracking

- 设计指尖级运动感知编码器,结构化融合本体感觉与触觉信息。
- 在序列追踪中定位误差降低15倍,优于传统融合方式。
- 无需显式监督,自动分配视觉负责位置、触觉负责旋转。
灵巧的手部操作需要持续的6维姿态追踪,但手指会遮挡相机视野。本文研究如何利用多指手上已有的稀疏触觉信号(包括本体感觉、近端力/力矩及二值接触),在遮挡条件下增强预训练的视觉姿态追踪器。提出一种运动感知的指尖级编码器,并通过三层次评估:帧级修正、序列开环追踪和闭环操作任务,系统比较其与四种替代设计的性能。实验表明:(i) 帧级评估无法区分编码器优劣,而序列追踪使架构差异放大至15倍;(ii) 结构化编码器学习到任务相关的跨模态门控机制,仅用视觉处理平移,一个注意力头专用于触觉处理旋转,且无显式监督;(iii) 采用4个令牌的紧凑指尖级表征优于扁平融合与关节级表示,后者因归一化主导抑制了视觉信息。验证显示追踪精度提升显著提高了下游重定向任务的成功率,并提供真实世界演示。项目页面见 https://cold-young.github.io/kine-fuse/。
原文摘要 · Abstract (English)
Dexterous in-hand manipulation requires continuous 6D pose tracking, yet the manipulating fingers inevitably occlude the object from the camera. We study how to structure the sparse haptic signals already available on multi-fingered hands, including proprioception, proximal force/torque, and binary contact, to complement a pretrained visual pose tracker under occlusion. We propose a kinematic-aware finger-level encoder and systematically compare it against four alternative designs through three levels of evaluation: per-frame refinement, sequential open-loop tracking, and closed-loop manipulation. Our experiments reveal that (i) per-frame evaluation cannot distinguish encoder quality, while sequential tracking amplifies architectural differences by up to 15 times; (ii) the structured encoder learns task-specific cross-modal gating, using vision exclusively for translation and dedicating one attention head to haptics for rotation, without explicit supervision; and (iii) compact finger-level tokenization with 4 tokens outperforms both flat fusion and joint-level representations, which suppress vision through norm dominance. We validate that improved tracking yields higher success in a downstream reorientation task and provide qualitative real-world demonstrations. Our project page is available at https://cold-young.github.io/kine-fuse/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。