arXiv:2608.15816cs.RO2026-08

让视觉语言模型学会用触觉调整动作,不改核心却大幅提效。

ViTaR: Visuo-Tactile Residual Adaptation for Foundation VLA Manipulation

论文配图:ViTaR: Visuo-Tactile Residual Adaptation for Foundation VLA Manipulation
图 1 · 摘自论文原文
  • 触觉不直接生成动作,而是作为修正量的调节器
  • 在7个高接触任务中成功率提升30.6个百分点至61.3%
  • 适合需要稳定物理交互的机器人操控场景

随着视觉-语言-动作(VLA)模型向真实场景部署发展,高接触操作暴露出一个关键盲区:这些策略虽具备广泛的视觉语义先验,却对局部接触事件无感,导致接触建立、丢失或失稳时输出相同动作。现有方法要么修改VLA内部结构,易引发灾难性遗忘;要么依赖近失败状态下的在线强化学习,使触觉对动作生成拥有无界影响,违背了使VLA具备泛化能力的先验假设。本文提出ViTaR,将触觉反馈从动作生成感知输入重构为执行调节器,仅在冻结的VLA基础上选择并缩放有界的残差修正,从构造上保留预训练能力。该方法分两阶段:效果引导建模通过结果导向的偏好证据判断局部修正是否合理;残差动作调节则根据实时多模态观测,将证据转化为连续可调增益的残差动作选择。在涵盖七项高接触任务的UniVTAC基准上,ViTaR平均成功率达61.3%,相比其冻结的VLA基线提升30.6个百分点,且优于专为触觉设计的基线。实体机器人实验验证了有界触觉调节能有效应对真实传感器噪声与动态差异。

原文摘要 · Abstract (English)

As Vision-Language-Action (VLA) models scale toward real-world deployment, contact-rich manipulation exposes a critical blind spot: these policies encode broad visual-semantic priors yet remain unaware of local contact events, producing identical actions whether contact is established, lost, or destabilized. Existing remedies either modify VLA internals, risking catastrophic forgetting, or demand online reinforcement under near-failure contact conditions. Both grant tactile unbounded influence over action generation, conflicting with the priors that make VLAs generalizable. We introduce ViTaR, which reframes tactile feedback from an action-generating perceptual input to an execution modulator that selects and scales bounded residual corrections atop a frozen VLA, preserving pretrained capabilities by construction. ViTaR decomposes adaptation into two stages: Effect-Guided Modeling determines whether and which correction is locally justified via outcome-grounded preference evidence, and Residual Action Modulation converts this evidence into a residual choice with continuously scaled gain from real-time visuotactile observations. On the UniVTAC benchmark spanning seven contact-rich tasks, ViTaR achieves 61.3% average success, a 30.6 percentage-point improvement over its frozen VLA base that also surpasses purpose-built tactile baselines. Physical-robot experiments confirm that bounded tactile modulation transfers to real sensor noise and dynamics.

视觉触觉动作调控机器人操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。