arXiv:2607.02840cs.RO2026-07被引 2

用触觉增强的世界模型,让机器人自己纠正操作失误

TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training

论文配图:TACO: TActile World Model as a Self-COrrector forScalable VLA Post-Training
图 1 · 摘自论文原文
  • 通过触觉感知识别失败前状态,生成修正动作序列
  • 实测成功率提升44%,优于无触觉适应的基线模型
  • 适合需要精细触觉交互的机器人任务开发者

视觉-语言-动作(VLA)模型在机器人操作中表现出良好泛化能力,但在依赖接触的任务中仍易因微小接触扰动导致不可恢复的失败,且仅靠视觉难以检测。此类失败为局部而非任务级语义错误,因此引入触觉反馈进行自我修正的后训练可高效提升恢复能力。但依赖人工标注的监督成本过高。现有工作尝试用世界模型生成模拟轨迹以改进策略,但仅依赖视觉的世界模型可能产生视觉合理却触觉不一致的轨迹。为此,本文提出TACO:一种触觉感知的世界模型驱动框架,用于接触密集型操作的可扩展VLA后训练。基于真实机器人轨迹,TACO采用“识别-想象-标注”循环:统一的进展-动作模型利用进展估计识别临近失败状态;视觉-触觉生成模型生成局部修正片段;进展-动作模型为其标注可执行的纠正动作。为将触觉纠正监督融入VLA后训练,TACO结合知识隔离式触觉适配与优势条件化训练,使策略能学习想象中的修正动作,同时不破坏预训练的视觉-语言先验。该框架可将真实失败转化为想象中的多模态修正,实现迭代优化。在真实接触密集型操作任务上的实验表明,TACO相较基础策略成功率达44%绝对提升,相比无知识隔离触觉适配的策略提升32%。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models have shown promising generalization in robotic manipulation, but they still struggle with contact-rich tasks, where minor contact perturbations can cause unrecoverable failures that are hard to detect from vision alone. Since these failures are localized rather than task-level semantic errors, tactile-aware corrective post-training offers an efficient way to improve recovery. However, scaling such supervision through human intervention is costly. Recent works have explored world models to synthesize imagined rollouts for policy improvement, but vision-only world models may produce visually plausible yet contact-inconsistent trajectories. We therefore introduce TACO, a tactile-aware world-model-driven framework for scalable VLA post-training in contact-rich manipulation. Given real robot rollouts, TACO follows a Recognize-Imagine-Label loop with a tactile-aware world model: a unified progress-action model recognizes failure-adjacent states using progress estimates, a visuo-tactile generation model imagines local correction segments, and the progress-action model labels them with executable corrective actions. To incorporate tactile corrective supervision into VLA post-training, TACO combines knowledge-insulated tactile adaptation with advantage-conditioned training, enabling the policy to learn from imagined corrections without degrading pretrained visual-language priors. These components enable TACO to convert real-world failures into imagined visuo-tactile corrections for iterative VLA post-training. Experiments on real-world contact-rich manipulation tasks show that TACO achieves 44% absolute success rate improvement over the base policy and 32% over the policy without knowledge-insulated tactile adaptation.

机器人操作触觉感知世界模型后训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。