无需标定即可统一估计物体与环境的接触点,提升机器人操作鲁棒性。
UNIC: Learning Unified Multimodal Extrinsic Contact Estimation
- 融合视觉、本体和触觉信息,用场景可利用性图统一建模接触
- 在未见物体上达到9.6毫米平均查姆费距离误差,支持动态视角
- 不依赖相机标定或预设接触类型,适合复杂未知环境
富含接触的操作需要可靠估计外在接触——即被抓物体与环境之间的交互,为规划、控制和策略学习提供关键上下文。然而现有方法常依赖预设接触类型、固定抓取配置或相机标定等限制性假设,难以泛化至新物体或非结构化环境。本文提出UNIC,一种无需先验知识或相机标定的统一多模态外在接触估计框架。UNIC直接编码相机帧中的视觉观测,并以全数据驱动方式融合本体感知与触觉模态。它引入基于场景可利用性图的统一接触表示,捕捉多样接触形态,并采用随机掩码的多模态融合机制,实现鲁棒的多模态表征学习。大量实验证明,UNIC表现稳定:在未见接触位置上达到9.6毫米平均查姆费距离误差,对未见物体表现良好,对缺失模态具有鲁棒性,且能适应动态相机视角。这些结果确立了外在接触估计作为接触丰富操作中实用而通用的能力。概览与硬件实验视频见https://youtu.be/xpMitkxN6Ls?si=7Vgj-aZ_P1wtnWZN
原文摘要 · Abstract (English)
Contact-rich manipulation requires reliable estimation of extrinsic contacts-the interactions between a grasped object and its environment which provide essential contextual information for planning, control, and policy learning. However, existing approaches often rely on restrictive assumptions, such as predefined contact types, fixed grasp configurations, or camera calibration, that hinder generalization to novel objects and deployment in unstructured environments. In this paper, we present UNIC, a unified multimodal framework for extrinsic contact estimation that operates without any prior knowledge or camera calibration. UNIC directly encodes visual observations in the camera frame and integrates them with proprioceptive and tactile modalities in a fully data-driven manner. It introduces a unified contact representation based on scene affordance maps that captures diverse contact formations and employs a multimodal fusion mechanism with random masking, enabling robust multimodal representation learning. Extensive experiments demonstrate that UNIC performs reliably. It achieves a 9.6 mm average Chamfer distance error on unseen contact locations, performs well on unseen objects, remains robust under missing modalities, and adapts to dynamic camera viewpoints. These results establish extrinsic contact estimation as a practical and versatile capability for contact-rich manipulation. The overview and hardware experiment videos are at https://youtu.be/xpMitkxN6Ls?si=7Vgj-aZ_P1wtnWZN
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。