arXiv:2512.14666cs.ROcs.CV2025-12被引 13

让视觉语言动作模型在运行时通过环境反馈持续自适应,减少对演示数据依赖。

EVOLVE-VLA: Test-Time Training from Environment Feedback for Vision-Language-Action Models

  • 用进度估计算法生成测试时的自主反馈信号
  • 长任务提升8.6%,单样本学习提升22.0%
  • 无需任务特定演示即可实现跨任务泛化

实现真正自适应的具身智能需要代理不仅模仿静态示范,更需通过环境互动持续改进,如同人类通过练习掌握技能。视觉-语言-动作(VLA)模型通过利用大语言模型推进了机器人操作,但仍受限于监督微调(SFT):每项任务需数百次示范、僵化记忆轨迹,且部署条件偏离训练时无法适应。我们提出EVOLVE-VLA,一种测试时训练框架,使VLA能通过环境互动以极少量或零任务特定示范持续适应。核心挑战是替代测试时不可用的原始奖励信号,我们通过一个学习的进度估计算法提供密集反馈,并设计两种机制‘驯服’这一固有噪声信号:(1) 累积进度估计机制平滑点状噪声估计,(2) 渐进式视野扩展策略实现政策逐步演化。EVOLVE-VLA显著提升性能:长任务+8.6%,1样本学习+22.0%,并实现跨任务泛化——在未见过的任务上达到20.8%成功率(纯SFT为0%)。定性分析揭示出演示中不存在的新兴能力,如错误恢复与新策略。该工作标志着向真正学习与自适应的VLA迈出关键一步,超越静态模仿,迈向持续自我改进。

原文摘要 · Abstract (English)

Achieving truly adaptive embodied intelligence requires agents that learn not just by imitating static demonstrations, but by continuously improving through environmental interaction, which is akin to how humans master skills through practice. Vision-Language-Action (VLA) models have advanced robotic manipulation by leveraging large language models, yet remain fundamentally limited by Supervised Finetuning (SFT): requiring hundreds of demonstrations per task, rigidly memorizing trajectories, and failing to adapt when deployment conditions deviate from training. We introduce EVOLVE-VLA, a test-time training framework enabling VLAs to continuously adapt through environment interaction with minimal or zero task-specific demonstrations. The key technical challenge is replacing oracle reward signals (unavailable at test time) with autonomous feedback. We address this through a learned progress estimator providing dense feedback, and critically, we design our framework to ``tame'' this inherently noisy signal via two mechanisms: (1) an accumulative progress estimation mechanism smoothing noisy point-wise estimates, and (2) a progressive horizon extension strategy enabling gradual policy evolution. EVOLVE-VLA achieves substantial gains: +8.6\% on long-horizon tasks, +22.0\% in 1-shot learning, and enables cross-task generalization -- achieving 20.8\% success on unseen tasks without task-specific demonstrations training (vs. 0\% for pure SFT). Qualitative analysis reveals emergent capabilities absent in demonstrations, including error recovery and novel strategies. This work represents a critical step toward VLAs that truly learn and adapt, moving beyond static imitation toward continuous self-improvements.

VLA自适应测试时训练强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。