让生成式视觉语言动作模型通过回忆成功经验提升部署稳定性。
Retrieve-then-Steer: Online Success Memory for Test-Time Adaptation of Generative VLAs

- 用成功轨迹构建长期记忆,推理时检索并筛选可靠动作片段。
- 在长程和多阶段任务中,成功率显著提升,闭合回路更稳定。
- 无需更新参数,轻量级非参数化适配,适合实际机器人持续运行场景。
视觉-语言-动作(VLA)模型在通用机器人操作中展现出巨大潜力,但在局部部署条件下闭环可靠性常下降。现有评估通常将测试回合视为独立的零样本试验,而现实中的机器人常在相同或缓慢变化环境中重复作业,成功的执行过程可提供环境验证的可靠行为证据。本文研究持续部署场景,探讨部分能力的冻结式VLA能否通过复用测试时的成功经验来提升可靠性。提出一种在线成功记忆引导的测试时自适应框架:部署期间,机器人将进度校准的成功观测-动作片段存入长期记忆;推理时,检索状态相关的动作片段,通过轨迹一致性过滤不一致候选,并聚合为精英动作先验。为将此先验融入动作生成,引入置信度自适应先验引导机制,将其注入流匹配动作采样器的中间状态,并根据检索置信度动态调整引导强度。该设计使冻结式VLA能利用环境特异性成功经验,同时保持观测条件下的生成精炼能力。这一‘检索-引导’机制实现了轻量、非参数化的测试时自适应,无需参数更新。仿真与真实世界实验均显示,在长时序与多阶段任务中,任务成功率与闭环稳定性显著提升。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models show strong potential for general-purpose robotic manipulation, yet their closed-loop reliability often degrades under local deployment conditions. Existing evaluations typically treat test episodes as independent zero-shot trials. However, real robots often operate repeatedly in the same or slowly changing environments, where successful executions provide environment-verified evidence of reliable behavior patterns. We study this persistent-deployment setting, asking whether a partially competent frozen VLA can improve its reliability by reusing its successful test-time experience. We propose an online success-memory guided test-time adaptation framework for generative VLAs. During deployment, the robot stores progress-calibrated successful observation-action segments in a long-term memory. At inference, it retrieves state-relevant action chunks, filters inconsistent candidates via trajectory-level consistency, and aggregates them into an elite action prior. To incorporate this prior into action generation, we introduce confidence-adaptive prior guidance, which injects the elite prior into an intermediate state of the flow-matching action sampler and adjusts the guidance strength based on retrieval confidence. This design allows the frozen VLA to exploit environment-specific successful experience while preserving observation-conditioned generative refinement. This retrieve-then-steer mechanism enables lightweight, non-parametric test-time adaptation without requiring parameter updates. Simulation and real-world experiments show improved task success and closed-loop stability, especially in long-horizon and multi-stage tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。