提出分层触觉世界动作模型,提升机器人高接触场景操作成功率。
HiTac-WAM: A Hierarchical Tactile World Action Model for Contact-Rich Robot Manipulation

- 构建分层触觉预测框架,包含接触状态、形变场和滑移风险三级结构。
- 实测触觉预测F1达0.921,滑移预测AUPRC提升60.4%,形变误差降低17.6%。
- 适用于高接触任务如插USB、擦黑板,真实机器人成功率从31.1%升至72.2%。
世界动作模型联合预测未来的视觉观测与动作,而现有触觉感知变体通常将未来触觉表示为图像或潜在流,未建模组织触觉状态的物理依赖关系。本文提出HiTac-WAM,一种分层触觉世界动作模型,可在执行前预测每个候选动作片段的一系列未来触觉状态。该预测分解为接触状态、三维形变场与滑移风险,按有向层次结构组织,下游阶段基于上游阶段的停梯度信号进行条件化。定向注意力掩码使触觉查询可访问各候选动作的视频-动作上下文,同时阻止视频与动作查询访问触觉标记。规划阶段中,通过触觉预测与任务进展估计对候选动作片段进行排序;执行阶段,保留选定的触觉预测作为参考,当预测与实际触觉状态持续偏差时触发修正重规划。在芯片抓取、黑板擦拭与USB插入任务中,基于分层预测的选择将真实机器人平均成功率从31.1%提升至61.1%,系统整体达到72.2%。在相同训练预算下,有向层次结构相较仅预测形变的模型减少17.6%的三维位移L2误差,相较仅预测滑移的模型提升60.4%的滑移AUPRC。触觉预测的平均接触F1为0.921。
原文摘要 · Abstract (English)
World action models jointly predict future visual observations and actions, whereas existing tactile-aware variants typically represent future touch as an image or latent stream without modeling the physical dependencies that organize tactile states hierarchically. We present HiTac-WAM, a hierarchical tactile world action model that forecasts a sequence of future tactile states for each candidate action chunk before execution. The forecast factorizes into contact state, a 3D deformation field, and slip risk, organized as a directed hierarchy in which each downstream stage is conditioned on stop-gradient signals from preceding stages. A directed attention mask allows tactile queries to attend to the video-action context of each candidate while preventing video and action queries from attending to tactile tokens. For planning, HiTac-WAM ranks candidate action chunks using tactile forecasts and task-progress estimates. For execution, the selected tactile forecast is retained as a reference; persistent discrepancies between predicted and observed tactile states trigger corrective replanning. HiTac-WAM achieves a mean contact F1 of 0.921; under matched training budgets, the directed hierarchy reduces 3D displacement L2 error by 17.6% relative to the deformation-only predictor and improves slip AUPRC by 60.4% relative to the slip-only predictor. Across chip grasping, blackboard erasing, and USB insertion, selection guided by the hierarchical forecasts increases the average real-robot success rate from 31.1% to 61.1%, while the full system attains 72.2%.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。