让机器人通过模拟学习触觉反馈,提升复杂操作成功率
TacCoRL: Integrating Tactile Feedback into VLA via Simulation

- 用模拟环境联合训练视觉-语言-动作策略,融合触觉输入
- 在4个双臂接触任务中,成功率从50%提升至72.5%
- 无需真实世界大量触觉数据,适合实际部署的机器人控制
视觉-语言-动作(VLA)模型为机器人操作提供了强大的视觉、语言和动作先验,但仅依赖视觉常无法获取接触任务所需的局部接触状态。我们提出TacCoRL,一个可扩展的框架,通过模拟与真实协同训练及基于仿真强化学习,将触觉反馈注入VLA策略,无需大规模触觉预训练或大量真实世界接触探索。核心思想不仅是将触觉作为输入,更是学习在演示中罕见且硬件收集危险的近失败状态下,如何调节动作响应。我们使用与真实对齐的模拟器作为闭环训练环境。混合模拟与真实轨迹首先在预训练策略中初始化触觉条件动作。随后,基于可验证任务奖励的强化学习利用模拟接触滚动优化策略,强化促成任务完成的触觉条件动作;同时,真实轨迹上的监督目标确保优化后的策略锚定于部署时的视觉、触觉和动作分布。最终策略可直接部署到真实机器人,无需特权模拟状态或在线真实强化学习。在四个双臂接触丰富任务中,最终的视觉-触觉策略平均成功率达72.5%,优于基线50.0%。更多视频与细节见https://tac-corl.github.io/
原文摘要 · Abstract (English)
Vision-language-action (VLA) models provide strong visual, language, and action priors for robot manipulation, but visual observations alone often miss the local contact state required for contact-rich tasks. We present TacCoRL, a scalable framework that injects Tactile feedback into VLA policies and improves them through sim-real Co-training and simulation-based reinforcement learning (RL), without requiring large-scale tactile pretraining or extensive real-world contact exploration. The key idea is not only adding touch as an input, but learning how contact readings should modulate action responses in near-failure states that are rare in demonstrations and risky to collect on hardware. We use a real-aligned simulator as a closed-loop training environment for contact interaction. Mixed simulated and real trajectories first warm-start tactile-conditioned actions in the pretrained policy. Reinforcement learning with verifiable task rewards then optimizes the policy using simulated contact rollouts. It reinforces tactile-conditioned actions that lead to task completion, while a supervised objective on real trajectories keeps the refined policy anchored to deployment visual, tactile, and action distributions. The resulting policy transfers directly to the real robot without privileged simulation state or online real-world RL. Across four bimanual contact-rich tasks, the final visuo-tactile policy achieves an average success rate of 72.5%, compared to baseline of 50.0%. Result videos and more details are available at https://tac-corl.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。