arXiv:2606.25800cs.LGcs.RO2026-06

让视觉语言动作模型在线适应更稳定,靠的是动作空间的自蒸馏。

ROAD-VLA: Robust Online Adaptation via Self-Distillation for Vision-Language-Action Models

论文配图:ROAD-VLA: Robust Online Adaptation via Self-Distillation for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 用动作词元的置信度扰动生成近邻教师,实现动作级密集监督。
  • 在7个机器人任务中,对分布内/外变化均优于PPO,提升显著。
  • 适合需要持续学习的机器人系统,尤其在奖励稀疏场景下。

视觉-语言-动作(VLA)模型的有效在线适应仍具挑战性,因稀疏奖励无法为高维自回归动作策略提供有效监督。尽管自蒸馏理论上可提供更密集的训练信号,但我们发现基于文本的优先教师(如演示、检索经验或高层计划)在VLA适应中效果不佳,暴露出符号引导与低层机器人动作之间的模态鸿沟。为此,我们提出ROAD-VLA,一种基于优势引导的自蒸馏框架:通过校准的优势估计扰动动作词元的对数概率,直接在动作空间构建近邻教师。该方法将稀疏奖励转化为密集的词元级监督,同时保持教师与当前策略的接近性。我们进一步推导出在优势校准和教师匹配准确时的策略改进下界。在包含分布内与分布外变化的7个机器人操作环境中,ROAD-VLA几乎在所有设置中均优于PPO,展现出鲁棒的在线适应能力。

原文摘要 · Abstract (English)

Effective online adaptation of vision-language-action (VLA) models remains challenging, as sparse rewards provide weak supervision for high-dimensional autoregressive action policies. Although self-distillation can in principle provide denser training signals, we find that text-based privileged teachers conditioned on demonstrations, retrieved experiences, or high-level plans are ineffective for VLA adaptation, exposing a modality gap between symbolic guidance and low-level robot actions. We propose ROAD-VLA, an advantage-guided self-distillation framework that constructs a proximal teacher directly in action space by perturbing action-token logits with calibrated advantage estimates. This converts sparse rewards into dense token-level supervision while keeping the teacher close to the current policy. We further derive a policy-improvement lower bound under calibrated advantages and accurate teacher matching. Across seven robotic manipulation environments with in-distribution and out-of-distribution shifts, ROADVLA outperforms PPO in nearly all settings, demonstrating robust online VLA adaptation.

机器人在线学习自蒸馏强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。