arXiv:2608.25585cs.RO2026-08被引 1

让机器人在不重新训练的情况下,快速适应新任务。

RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation

论文配图:RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation
图 1 · 摘自论文原文
  • 用语义对齐的检索机制,从经验中找合适动作
  • 在虚拟和真实机械臂上成功率达90%以上
  • 适合需要快速部署的智能机器人场景

视觉-语言-动作(VLA)模型为通用机器人操作提供了基础,但在面对新任务分布时仍显脆弱。尽管上下文模仿学习(ICIL)无需训练即可适应,但现有方法受限于浅层检索机制和行为惯性,难以将专家经验转化为有效动作。为此,我们提出RA-VLA,一种融合行为对齐检索与可落地执行流程的增强型VLA框架。通过在可扩展架构中严格遵循功能线索,该方法实现无缝任务迁移,同时保持高效推理。在LIBERO基准和真实UR5e环境中验证表明,RA-VLA在成功率与计算效率方面均优于现有方法,构建了一个高效的免训练机器人自适应框架。

原文摘要 · Abstract (English)

Vision-Language-Action (VLA) models provide a versatile foundation for general robotic manipulation, yet they exhibit significant brittleness when confronted with novel task distributions. While In-Context Imitation Learning (ICIL) offers a training-free alternative, existing frameworks suffer from an adaptation bottleneck that hinders the effective translation of expert context to executable actions. This failure originates from superficial retrieval mechanisms and an inherent behavioral inertia that anchors the policy to its pre-trained priors. To address these limitations, we present RA-VLA, a retrieval-augmented VLA framework that integrates behavior-aligned context retrieval with a grounded execution pipeline. By enforcing faithful adherence to functional cues within a scalable architecture, RA-VLA facilitates seamless task adaptation while preserving inference efficiency. Our empirical evaluations across the LIBERO benchmark and a real-world UR5e environment demonstrate that RA-VLA achieves superior success rates and computational efficiency, establishing a robust framework for training-free robotic adaptation.

机器人自适应VLA

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。