零样本融合视觉与触觉信息,实现机器人抓取中6D物体姿态精准估计。
ViTa-Zero: Zero-shot Visuotactile Object 6D Pose Estimation
- 用视觉模型做主干,结合触觉与本体感知的物理约束进行测试时优化。
- 在真实机器人上测试,平均ADD-S AUC提升55%,位置误差降低80%。
- 无需训练数据即可跨场景通用,适合实际抓取与双臂协作任务。
6D物体姿态估计是机器人操作中的关键挑战。尽管视觉与触觉融合方法已展现潜力,但受限于触觉数据稀缺,泛化能力差。本文提出ViTa-Zero,一种零样本视觉触觉姿态估计框架。核心创新在于以视觉模型为骨干,基于触觉与本体感知的物理约束进行可行性检验和测试时优化。具体地,将夹爪-物体交互建模为弹簧-质量系统,触觉传感器产生吸引力,本体感知生成排斥力。在真实机器人平台上的实验验证了该框架的有效性,适用于多种视觉骨干和操作场景,包括抓取、拾取和双臂交接。相比纯视觉模型,本方法克服了若干严重失败模式,并能持续追踪手内物体姿态。实验表明,平均ADD-S AUC提升55%,ADD提升60%,位置误差较FoundationPose降低80%。
原文摘要 · Abstract (English)
Object 6D pose estimation is a critical challenge in robotics, particularly for manipulation tasks. While prior research combining visual and tactile (visuotactile) information has shown promise, these approaches often struggle with generalization due to the limited availability of visuotactile data. In this paper, we introduce ViTa-Zero, a zero-shot visuotactile pose estimation framework. Our key innovation lies in leveraging a visual model as its backbone and performing feasibility checking and test-time optimization based on physical constraints derived from tactile and proprioceptive observations. Specifically, we model the gripper-object interaction as a spring-mass system, where tactile sensors induce attractive forces, and proprioception generates repulsive forces. We validate our framework through experiments on a real-world robot setup, demonstrating its effectiveness across representative visual backbones and manipulation scenarios, including grasping, object picking, and bimanual handover. Compared to the visual models, our approach overcomes some drastic failure modes while tracking the in-hand object pose. In our experiments, our approach shows an average increase of 55% in AUC of ADD-S and 60% in ADD, along with an 80% lower position error compared to FoundationPose.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。