arXiv:2607.26991cs.RO2026-07

让视觉语言动作模型在测试时自适应地调整行为,提升复杂任务成功率。

RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models

论文配图:RL$^2$-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 用强化学习优化视觉语言动作模型的潜在空间,生成多样化动作。
  • 在可能失败时才激活干预,成功时保持原动作,避免干扰。
  • 实测在仿真和真实世界中均显著提升复杂任务成功率。

尽管视觉语言动作(VLA)模型具备出色的视觉运动能力,但在挑战性及域外任务上性能常下降。现有测试时调优方法虽无需额外训练,但动作样本仍集中于相似行为,继承相关失败模式。且所有时刻采用相同干预策略,未考虑基础策略的成功概率。为此,我们提出RL²,一种基于强化学习的自适应推理时调优框架。首先,在从VLA动作专家提取的丰富潜在表示上训练轻量级离线强化学习策略,并在推理时将该策略的流动速度与冻结的VLA组合。此组合策略融合大规模模仿学习的行为先验与离线强化学习带来的超越主导演示模式的动作多样性。进一步发现,推理时调优遵循不同的缩放规律:当基础VLA可能失败时,动作多样性最有益;而成功概率高时,扰动反而有害。基于此,RL²仅在预测失败时激活组合调优。在SIMPLER和PolaRiS基准测试中,RL²在域外设置下成功率达+17.3%提升,消融实验与缩放研究验证了潜在表示与强化学习训练的重要性。最终,真实世界实验表明这些增益可跨仿真迁移,确立了RL²作为实际部署中模块化调优框架的可行性。

原文摘要 · Abstract (English)

Despite the impressive visuomotor capabilities enabled by Vision-Language-Action (VLA) models, their performance often degrades on challenging and out-of-domain tasks. Recent test-time steering and scaling methods improve performance without extensive data collection and retraining, but action samples often remain concentrated around similar behaviors and therefore inherit correlated failure modes. Moreover, existing methods apply the same intervention strategy at every timestep, regardless of whether the base policy is already likely to succeed. To address these limitations, we introduce $RL^2$, an adaptive inference-time steering framework that leverages Reinforcement Learning on VLA Latents. First, we train a lightweight offline RL policy conditioned on expressive latents extracted from the VLA action expert and compose its flow velocity with that of the frozen VLA during inference. This compositional steering strategy combines the behavioral priors of large-scale imitation learning with the action diversity induced by offline RL beyond dominant demonstration modes. We further discover that inference-time steering follows fundamentally different scaling laws under success and failure states, revealing that action diversity is most beneficial when the base VLA is likely to fail, but can unnecessarily perturb already-accurate actions when success is likely. Building on this insight, $RL^2$ activates compositional steering only when failure is predicted. Across the SIMPLER and PolaRiS benchmarks, $RL^2$ improves success rates by up to +17.3% in out-of-domain settings, while ablations and scaling studies demonstrate the importance of latent representations and RL training. Finally, real-world experiments demonstrate that these gains transfer beyond simulation, establishing $RL^2$ as a practical and modular steering framework for VLA deployment.

VLA强化学习推理调优动作多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。