arXiv:2601.06748cs.RO2026-01ACL被引 11

让机器人在运行时自主调整策略,应对未知环境。

On-the-Fly VLA Adaptation via Test-Time Reinforcement Learning

  • 运行时用强化学习动态优化动作策略,不需重新训练。
  • 在模拟和真实场景中提升任务成功率与稳定性。
  • 适合需要自适应能力的机器人部署场景。

视觉-语言-动作模型(VLA)已成为通用机器人学习的强大范式,可将视觉观测和自然语言指令映射为可执行动作。然而,现有方法主要依赖监督微调或训练时强化学习,需显式微调阶段、人工干预或受控数据收集,难以适用于模拟或真实世界中需自主响应动态环境的部署。为此,我们提出测试时强化学习框架TT-VLA,实现推理过程中的即时策略适配。该框架设计密集奖励机制,利用逐步任务进展信号在测试时优化动作策略,同时保留原有监督微调/强化学习训练的先验知识,有效补充现有VLA模型。实验证明,该方法显著提升动态、未见场景下的适应性、稳定性和任务成功率,涵盖模拟与真实世界设置。我们认为,TT-VLA为构建可自我改进、可部署的VLA迈出了关键一步。

原文摘要 · Abstract (English)

Vision-Language-Action models have recently emerged as a powerful paradigm for general-purpose robot learning, enabling agents to map visual observations and natural-language instructions into executable robotic actions. Though popular, they are primarily trained via supervised fine-tuning or training-time reinforcement learning, requiring explicit fine-tuning phases, human interventions, or controlled data collection. Consequently, existing methods remain unsuitable for challenging simulated- or physical-world deployments, where robots must respond autonomously and flexibly to evolving environments. To address this limitation, we introduce a Test-Time Reinforcement Learning for VLAs (TT-VLA), a framework that enables on-the-fly policy adaptation during inference. TT-VLA formulates a dense reward mechanism that leverages step-by-step task-progress signals to refine action policies during test time while preserving the SFT/RL-trained priors, making it an effective supplement to current VLA models. Empirical results show that our approach enhances overall adaptability, stability, and task success in dynamic, previously unseen scenarios under simulated and real-world settings. We believe TT-VLA offers a principled step toward self-improving, deployment-ready VLAs.

机器人学习强化学习自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。