用视觉和触觉联合调整机器人执行动作,提升复杂操作成功率。
Inference-time Policy Steering via Vision and Touch

- 分层优化:视觉选大方向,触觉微调局部动作
- 真实任务中成功率提升51%,优于单模态33%以上
- 首创触觉奖励直接在隐空间评分,支持语义级控制
推理时调整通过部署阶段验证候选动作来适应预训练生成式机器人策略。以往方法仅依赖视觉观察,但在涉及接触的操控任务中,成功不仅取决于全局任务进展,还依赖于接触力等细微局部交互。我们提出ViTaL,一种基于视觉与触觉的推理时调整框架,将多模态引导建模为双层优化问题。高层采用视觉采样与验证进行长时程行为选择,决定应执行的动作模式;低层则通过触觉引导的扩散编辑,在较短时域内优化选定的动作序列以满足局部接触需求。为支持结果导向的调整,ViTaL学习了一个视觉-触觉联合隐空间模型,并使用语义对齐的视觉与触觉验证器,包括一种新颖的文本条件触觉奖励,可直接在隐空间中评分预测的触觉未来。在三个真实世界的高接触密度操控任务中,相比基础策略,ViTaL整体成功率提升51%,优于单模态调整至少33%,超过朴素多模态融合至少20%。
原文摘要 · Abstract (English)
Inference-time steering adapts pre-trained generative robot policies during deployment by verifying candidate actions before execution. While prior methods typically perform this verification only with visual observations, vision alone is often insufficient for contact-rich manipulation, where success depends on both global task progress and subtle local interactions such as contact force. We introduce ViTaL, a visuo-tactile inference-time steering framework that formulates multimodal guidance as a bi-level optimization problem. At the high level, visual sampling-and-verification performs long-horizon mode selection, deciding what behavior the robot should execute. At the low level, tactile-guided diffusion editing refines the selected action sequence over a shorter horizon to satisfy local contact requirements. To support outcome-based steering, ViTaL learns a visuo-tactile latent world model and employs semantically aligned visual and tactile verifiers, including a novel text-conditioned tactile reward that scores predicted tactile futures directly in latent space. Across three real-world contact-rich manipulation tasks, ViTaL improves overall success by 51% over the base policy, outperforms unimodal steering by at least 33%, and exceeds naive multimodal fusion by at least 20%. Website: https://yilin-wu98.github.io/vital_website.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。