融合视觉与触觉信息,实现更精准的6维物体姿态实时追踪。
V-HOP: Visuo-Haptic 6D Object Pose Tracking
- 设计统一触觉表征,支持多种夹持器和传感器形态。
- 在真实场景中显著优于纯视觉追踪方法,尤其在新设备上表现稳定。
- 适合机器人抓取、精密操作等需要多模态感知的任务。
人类在操作物体时自然融合视觉与触觉信息以实现鲁棒感知,任一模态缺失都会导致性能下降。受此启发,已有研究尝试结合视觉与触觉反馈进行物体位姿估计。尽管这些方法在受控环境或合成数据集上表现良好,但在真实场景中常因对不同夹持器、传感器布局或仿真到现实的泛化能力差而逊于纯视觉方法。此外,它们通常逐帧独立估计位姿,导致序列跟踪不连贯。为此,我们提出一种新型统一触觉表征,有效应对多种夹持器形态。基于该表征,构建基于视觉-触觉变换器的物体位姿追踪框架,无缝融合视觉与触觉输入。我们在自建数据集及Feelsight数据集上验证了该框架,结果表明其在复杂序列上取得显著提升。特别地,该方法在新夹持器、新物体和不同传感器类型(包括基于触点和基于视觉的触觉传感器)下均表现出优异泛化能力。真实实验显示,其性能远超现有先进视觉追踪方法。进一步将实时位姿追踪结果嵌入运动规划,成功实现精确操控任务,凸显多模态感知的优势。
原文摘要 · Abstract (English)
Humans naturally integrate vision and haptics for robust object perception during manipulation. The loss of either modality significantly degrades performance. Inspired by this multisensory integration, prior object pose estimation research has attempted to combine visual and haptic/tactile feedback. Although these works demonstrate improvements in controlled environments or synthetic datasets, they often underperform vision-only approaches in real-world settings due to poor generalization across diverse grippers, sensor layouts, or sim-to-real environments. Furthermore, they typically estimate the object pose for each frame independently, resulting in less coherent tracking over sequences in real-world deployments. To address these limitations, we introduce a novel unified haptic representation that effectively handles multiple gripper embodiments. Building on this representation, we introduce a new visuo-haptic transformer-based object pose tracker that seamlessly integrates visual and haptic input. We validate our framework in our dataset and the Feelsight dataset, demonstrating significant performance improvement on challenging sequences. Notably, our method achieves superior generalization and robustness across novel embodiments, objects, and sensor types (both taxel-based and vision-based tactile sensors). In real-world experiments, we demonstrate that our approach outperforms state-of-the-art visual trackers by a large margin. We further show that we can achieve precise manipulation tasks by incorporating our real-time object tracking result into motion plans, underscoring the advantages of visuo-haptic perception. Project website: https://ivl.cs.brown.edu/research/v-hop
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。