arXiv:2606.29892cs.ROcs.AI2026-06

让视觉语言动作模型靠自信度自我优化,无需外部奖励。

Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models

论文配图:Trust Your Instincts: Confidence-Driven Test-Time RL for Vision-Language-Action Models
图 1 · 摘自论文原文
  • 用高置信度轨迹作为内在奖励信号,实现测试时自增强学习。
  • 在LIBERO和RoboTwin上超越监督基线,接近有真实奖励时的最优性能。
  • 适配多种模型架构,动态平衡探索与训练稳定性,适合无奖励场景

强化学习(RL)已成为推动视觉-语言-动作模型(VLAs)突破静态模仿学习的关键。然而,现有方法通常依赖外部环境反馈,通过预设的成功信号指导策略更新。本文发现,离散动作的VLAs具备内部评估能力:生成置信度越高的轨迹,成功概率显著更高。基于此,我们提出T^2VLA(测试时VLA),一种不依赖特定架构的测试时强化学习框架,使VLAs实现自举式策略改进。该方法不使用外部奖励,而是利用轨迹与高置信度专家示范之间的相似性作为内在奖励信号。同时,提出置信度驱动的双专家自举机制,动态平衡局部伪专家以促进探索、全局专家池以保障训练稳定。在LIBERO和RoboTwin基准上的大量实验表明,T^2VLA持续优于监督基线,并在无外部奖励情况下逼近具有真实奖励的极值性能。此外,T^2VLA可适应多种VLAs范式,包括OpenVLA-OFT与pi系列模型。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has become indispensable for pushing Vision-Language-Action Models (VLAs) beyond static imitation learning. However, existing RL methods typically require external environmental feedback, relying on predefined success signals to guide policy updates. In this work, we show that VLA models possess useful internal evaluative capabilities: in discrete-action VLAs, trajectories with higher generation confidence are significantly more likely to succeed. Based on this observation, we introduce T^2VLA (Test-time VLA), an architecture-agnostic test-time RL framework that enables VLA models to achieve self-bootstrapping policy improvement. Instead of relying on external rewards, T^2VLA leverages trajectory-level similarity to high-confidence expert demonstrations as an intrinsic reward signal. In addition, we propose a Confidence-Driven Dual Expert Bootstrapping mechanism, which dynamically balances a Local Pseudo-Expert for exploration and a Global Expert Pool for training stability. Extensive experiments on the LIBERO and RoboTwin benchmarks show that T^2VLA consistently outperforms supervised baselines and approaches oracle RL performance with ground-truth rewards, achieving effective improvement without external reward feedback. Furthermore, T^2VLA adapts to distinct VLA paradigms, including both OpenVLA-OFT and the pi series.

强化学习视觉语言自举学习无奖励

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。