arXiv:2606.09337cs.RO2026-06被引 1

让机器人通过触觉反馈在线优化抓取动作,提升复杂操作成功率。

TORL-VLA: Tactile Guided Online Reinforcement Learning for Contact-Rich Manipulation

论文配图:TORL-VLA: Tactile Guided Online Reinforcement Learning for Contact-Rich Manipulation
图 1 · 摘自论文原文
  • 用触觉信号引导策略生成参考动作并预测受力变化
  • 在线强化学习模块使策略在真实场景中持续优化
  • 适合需要精细触觉控制的长周期操作任务

视觉-语言-动作(VLA)模型已成为机器人操作的重要框架,近期研究将触觉或力反馈引入VLA以应对接触密集型任务。然而,这些模型通常作为离线策略部署,当接触条件偏离训练分布时无法进行在线适应,导致接触力不当和重试效率低下。为此,我们提出TORL-VLA,一种基于触觉引导的在线强化学习框架,将触觉反馈与策略精炼相结合,用于接触密集型操作。该方法引入一种由触觉驱动的力矩感知VLA,用于预测参考动作和未来力矩序列,同时采用轻量级在线强化学习模块对参考动作进行优化。为稳定混合探索性策略生成数据与人工干预数据的学习,我们设计了干预屏蔽评价网络,防止干预后的成功被错误归因于干预前的策略行为。在长周期接触密集型任务(包括搭扣操作、咖啡杯放置和鸡蛋处理)的真实机器人实验中,TORL-VLA在子任务与完整任务层面均显著提升成功率,并在时间约束条件下实现更高的执行效率,优于多个强基线模型。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have become a powerful framework for robotic manipulation, and recent studies have introduced tactile or force feedback into VLAs to address contact-rich tasks. However, these models are typically deployed as offline policies. When contact conditions shift from the training distribution, the policy cannot perform online adaptation, leading to problems such as inappropriate contact forces and inefficient retries. Therefore, we propose TORL-VLA, a tactile-guided online reinforcement learning framework that couples tactile feedback with policy refinement for contact-rich manipulation. Our method introduces a tactile-derived wrench-aware VLA to predict reference actions and future wrench sequences, while a lightweight online RL module is used to refine the reference actions. To stabilize learning from mixed exploratory policy-generated and human-intervention data, we introduce an intervention-censored critic that prevents post-intervention success from being wrongly credited to policy-generated actions preceding intervention. Real-robot experiments on long-horizon contact-rich tasks, including latch manipulation, coffee-cup placement, and egg handling, show that TORL-VLA improves success rates at both subtask and full-task levels, as well as time-bounded execution efficiency over strong baselines. Project page: https://torl-vla.github.io/

机器人操控触觉反馈在线学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。