让机器人边执行边用触觉反馈调整动作,提升操作精度
TacForcing: Streaming Action Generation with Execution-Time Tactile Feedback

- 用实时触觉信号动态生成动作,替代原有静态预测
- 模拟与真实任务中成功率分别达65%和69%,优于基线
- 仅需单一模型,避免额外控制器,降低系统复杂度
接触丰富的操作需要适应在动作期内可能发生显著变化的接触状态。然而,基于块的视觉-语言-动作模型在执行前就依据观测预测完整动作块,导致执行期间触觉条件滞后。现有触觉响应方法通常依赖独立的高频控制器,增加架构与训练复杂性。本文提出TacForcing,一种流式动作生成框架,有效融合执行时触觉反馈。不使用独立反应控制器,而是将标准动作专家替换为流式动作专家,根据执行过程中获取的动态触觉观测生成动作。TacForcing还引入执行感知触觉注意力(EATA),仅对临近执行的动作进行触觉条件化,减少触觉获取与动作执行之间的时序错配。在六个模拟UniVTAC任务和三个真实世界接触丰富操作任务中,TacForcing分别实现平均成功率65%和69%,在两类场景下均优于强基线。
原文摘要 · Abstract (English)
Contact-rich manipulation requires adapting to contact states that can evolve substantially within an action horizon. However, chunk-based vision-language-action models predict complete action chunks from observations collected before execution, leaving tactile conditioning stale during execution. Existing tactile-reactive approaches typically rely on separate high-frequency controllers, which increase both architectural and training complexity. In this paper, we introduce TacForcing, a streaming action-generation framework that effectively incorporates execution-time tactile feedback. Instead of employing a separate reactive controller, TacForcing replaces the standard action expert with a streaming action expert to generate actions conditioned on the evolving tactile observations acquired during execution. TacForcing also introduces Execution-Aware Tactile Attention (EATA), which restricts tactile conditioning to actions nearing execution, thereby reducing the temporal mismatch between tactile acquisition and action execution. Across six simulated UniVTAC tasks and three real-world contact-rich manipulation tasks, TacForcing achieves average success rates of 65% and 69%, respectively, outperforming strong baselines in both settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。