arXiv:2607.07287cs.RO2026-07被引 3

让机器人同时预判和应对触觉反馈,提升灵巧操作的稳定性和成功率。

TouchWorld: A Predictive and Reactive Tactile Foundation Model for Dexterous Manipulation

论文配图:TouchWorld: A Predictive and Reactive Tactile Foundation Model for Dexterous Manipulation
图 1 · 摘自论文原文
  • 分层策略:先规划任务,再生成动作,最后高频修正触觉误差。
  • 在6个高难度任务中,干净环境下成功率达65.0%,扰动下仍达53.7%。
  • 适合需要精准触觉控制的机器人操作场景,如装配、抓取等。

日常环境中灵巧操作需兼具预测与应变能力:机器人须预判接触演化,同时快速纠正因滑移、对不准、抓握不稳或力不匹配引发的局部误差。视觉与语言提供语义和几何指导,但无法可靠揭示力、滑移、接触稳定性等隐藏接触状态。尽管触觉感知可暴露这些物理线索,现有策略多将触觉视为单一流程中的低频观测,将慢速任务推理、动作生成与快速接触反馈耦合于同一循环。本文提出TouchWorld,一种用于灵巧操作的预测-响应式触觉基础模型。其采用分层策略,分离视觉-语言子任务规划、触觉世界模型预测、视触觉目标条件动作生成以及高频触觉残差优化。高层规划层生成可执行子任务并预测触觉目标;视触觉目标条件策略生成名义动作片段;触觉条件优化策略利用近期触觉与本体感知反馈进行在线残差修正。通过将触觉用作预测参考与快速反馈信号,TouchWorld在保持视觉-语言-动作策略语义泛化性的同时,显著提升局部接触适应能力。在六个长时程、高接触密度的灵巧操作任务中,该模型在干净设置下成功率达65.0%,在人类干扰下为53.7%,分别优于最强基线15.7和18.5个百分点。

原文摘要 · Abstract (English)

Dexterous manipulation in everyday environments requires both anticipation and reaction: a robot must predict how contact should evolve while rapidly correcting local errors caused by slip, misalignment, unstable grasping, or force mismatch. Vision and language provide semantic and geometric guidance, but they cannot reliably reveal hidden contact states such as force, slip, and contact stability. Although tactile sensing exposes these physical cues, most existing policies treat touch as a low-frequency observation stream within a monolithic action model, coupling slow task reasoning, action generation, and fast contact feedback in a single loop. We introduce TouchWorld, a predictive-and-reactive tactile foundation model for dexterous manipulation. TouchWorld uses a hierarchical policy that separates vision-language subtask planning, tactile world-model prediction, visuo-tactile goal-conditioned action generation, and high-frequency tactile residual refinement. A High-Level Planning Layer produces executable subtasks and predicts tactile subgoals; a Visuo-Tactile Goal-Conditioned Policy generates nominal action chunks; and a Tactile-Conditioned Refinement Policy performs online residual correction using recent tactile and proprioceptive feedback. By using touch as both a predictive contact reference and a fast feedback signal, TouchWorld preserves the semantic generalization of vision-language-action policies while improving local contact adaptation. Across six long-horizon and contact-rich dexterous manipulation tasks, TouchWorld achieves 65.0% success in the clean setting and 53.7% success under human perturbations, outperforming the strongest baseline by 15.7 and 18.5 percentage points, respectively.

触觉模型灵巧操作分层控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。