arXiv:2505.08213cs.RO2025-05被引 1

用手腕摄像头和惯性传感器融合实现高精度手部自感知,无漂移、泛化强。

HandCept: A Visual-Inertial Fusion Framework for Accurate Proprioception in Dexterous Hands

  • 用视觉与惯性数据实时融合,零样本学习估计关节角度。
  • 误差2°~4°,无漂移,优于纯视觉或纯惯性方法。
  • 开源高保真渲染管线,支持从仿真到现实的迁移训练。

随着机器人向通用操作发展,灵巧手的重要性日益凸显。然而,受限于体积与通用性,灵巧手的本体感知仍是一大瓶颈。本文提出HandCept,首个面向灵巧手的视觉-惯性本体感知框架,解决了动态环境下视觉与惯性测量噪声和漂移导致的关节角度估计难题。该框架利用腕戴式RGB-D相机与9轴IMU,通过无延迟扩展卡尔曼滤波器实现实时融合,采用零样本学习策略。实验表明,其关节角度估计误差普遍在2°至4°之间,且无可观测漂移,显著优于纯视觉或纯惯性方法。同时验证了IMU系统稳定性和统一基座帧对校准的简化作用。为支持仿真到现实的迁移,本文开源了高保真渲染管道,对无需真实标注数据的训练至关重要。本工作为灵巧手本体感知提供了鲁棒、可泛化的解决方案,对机器人操作与人机交互具有重要意义。

原文摘要 · Abstract (English)

As robotics progresses toward general manipulation, dexterous hands are becoming increasingly critical. However, proprioception in dexterous hands remains a bottleneck due to limitations in volume and generality. In this work, we present HandCept, the first visual-inertial proprioception framework designed to overcome the challenges of traditional joint angle estimation methods for dexterous hands. HandCept addresses the difficulty of achieving accurate and robust joint angle estimation in dynamic environments where both visual and inertial measurements are prone to noise and drift. It leverages a zero-shot learning approach using a wrist-mounted RGB-D camera and 9-axis IMUs, fused in real time via a latency-free Extended Kalman Filter (EKF). Our results show that HandCept achieves joint angle estimation errors generally between $2^{\circ}$ and $4^{\circ}$ without observable drift, outperforming visual-only and inertial-only methods. Furthermore, we validate the stability and uniformity of the IMU system, demonstrating that a common base frame across IMUs simplifies system calibration. To support sim-to-real transfer, we also open-source our high-fidelity rendering pipeline, which is essential for training without real-world ground truth. This work offers a robust, generalizable solution for proprioception in dexterous hands, with significant implications for robotic manipulation and human-robot interaction. https://github.com/huangjund/blenderYCB

灵巧手本体感知视觉惯性融合机器人操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。