用本体感知对齐视觉与力觉,实现零样本高精度装配
Zero-Shot Sim-to-Real Contact-Rich Assembly via Proprioception-Anchored Cross-Modal Pretraining

- 以关节状态预测为监督,对齐仿真与真实环境的多模态特征
- 四类装配任务平均成功率93.3%,模拟到现实性能下降仅2.7个百分点
- 无需微调或姿态追踪,对扰动鲁棒性强,适合工业装配场景
接触丰富的装配任务因需亚毫米级空间精度和持续接触下的力信息可靠解析而极具挑战。尽管基于仿真的强化学习提供了可扩展的训练范式,但视觉观测、接触动力学及力/扭矩(F/T)测量在仿真与真实间存在差异,常限制策略迁移。我们观察到,本体感知在域间相对一致:校准后的关节位置与一致计算的关节速度在仿真与硬件中高度吻合。基于此,我们提出PACE(本体感知锚定的跨模态编码器),通过预测本体状态转移来监督时序视觉与F/T表示。静态域特定因素(如光照、纹理、传感器偏移)对关节运动信息贡献极小,该目标促使编码器抑制这些无关因素,保留任务相关的运动线索。在冻结的PACE特征上训练的策略直接部署于硬件,无需真实世界微调或物体姿态追踪。在四项接触丰富装配任务中,PACE实现平均93.3%的真实世界成功率,模拟到现实性能下降仅2.7个百分点,且对显著降低基于姿态和学习融合基线性能的扰动仍保持鲁棒性。
原文摘要 · Abstract (English)
Contact-rich assembly remains challenging because it requires submillimeter spatial accuracy and reliable interpretation of forces during sustained contact. Although simulation-based reinforcement learning offers a scalable training paradigm, discrepancies in visual observations, contact dynamics, and force/torque (F/T) measurements often limit policy transfer. We observe that proprioception is comparatively consistent across domains because calibrated joint positions and consistently computed joint velocities align closely between simulation and hardware. Based on this observation, we present PACE (Proprioception-Anchored Cross-Modal Encoder), which supervises temporal visual and F/T representations by predicting proprioceptive state transitions. Static domain-specific factors, including lighting, texture, and sensor bias, contain little information about joint motion; the proposed objective therefore encourages the encoder to suppress these factors while retaining task-relevant motion cues. Policies trained on frozen PACE features are deployed on hardware without real-world fine-tuning or object-pose tracking. Across four contact-rich assembly tasks, PACE attains an average real-world success rate of 93.3\% and only a 2.7-percentage-point sim-to-real drop, meanwhile remaining robust to perturbations that substantially degrade pose-based and learned-fusion baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。