arXiv:2510.14930cs.ROcs.LG2025-10中稿 · CoRL被引 25

用视觉触觉反馈提升机器人双手装配精度

VT-Refine: Learning Bimanual Assembly with Visuo-Tactile Feedback via Simulation Fine-Tuning

  • 结合真实演示与高保真触觉仿真,训练扩散策略
  • 在模拟环境中通过强化学习提升策略鲁棒性,成功率显著提高
  • 适合需要精细接触操作的工业装配场景

人类能凭借丰富的触觉反馈完成双手装配任务,而机器人仅靠行为克隆难以复现这一能力,因人类示范数据存在次优且多样性不足的问题。本文提出VT-Refine框架,融合真实示范、高保真触觉仿真与强化学习,实现精确、接触密集的双臂装配。首先在少量示范数据上使用同步视觉与触觉输入训练扩散策略;随后将该策略迁移至配备模拟触觉传感器的数字孪生环境,通过大规模强化学习进一步优化以增强鲁棒性和泛化能力。为实现准确的仿真到现实迁移,采用高分辨率压阻式触觉传感器获取法向力信号,并利用GPU加速仿真实现逼真建模。实验表明,VT-Refine在仿真和真实世界中均提升了装配性能,通过增加数据多样性并支持更有效的策略微调。

原文摘要 · Abstract (English)

Humans excel at bimanual assembly tasks by adapting to rich tactile feedback -- a capability that remains difficult to replicate in robots through behavioral cloning alone, due to the suboptimality and limited diversity of human demonstrations. In this work, we present VT-Refine, a visuo-tactile policy learning framework that combines real-world demonstrations, high-fidelity tactile simulation, and reinforcement learning to tackle precise, contact-rich bimanual assembly. We begin by training a diffusion policy on a small set of demonstrations using synchronized visual and tactile inputs. This policy is then transferred to a simulated digital twin equipped with simulated tactile sensors and further refined via large-scale reinforcement learning to enhance robustness and generalization. To enable accurate sim-to-real transfer, we leverage high-resolution piezoresistive tactile sensors that provide normal force signals and can be realistically modeled in parallel using GPU-accelerated simulation. Experimental results show that VT-Refine improves assembly performance in both simulation and the real world by increasing data diversity and enabling more effective policy fine-tuning. Our project page is available at https://binghao-huang.github.io/vt_refine/.

双臂装配触觉反馈仿真迁移扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。