用结构化示范让视觉语言动作模型实时自适应新任务,无需重新训练。
StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models

- 通过零成本转换轨迹为任务计划、子目标和3D运动描述,实现结构化示范。
- 在VLA-Arena上得分0.63,LIBERO成功率98.8%,超越现有模型。
- 支持人、机器人、虚拟现实多种示范输入,跨实体迁移能力强。
视觉语言动作(VLA)模型能理解指令并操作物体,但其性能在分布外(OOD)场景中常严重下降。传统方法需收集新数据并微调。本文提出StellaVLA,一种在推理时通过单个检索到的示范进行自适应的框架。核心思想是超越模仿专家行为,转而传达行为原因:一个自动化离线流程将原始轨迹转化为结构化示范,如任务计划、子目标描述和显式3D运动,且无需人工标注。作为上下文引导,该结构化示范使策略能推理任务而非复制像素轨迹,提升跨实体迁移能力(真实机器人、人手或XR示范)。采用并行双训练设计,在训练中通过联合动作与语言目标内化此推理机制;推理时仅使用动作专家,保持实时高频控制且无延迟增加。在VLA-Arena排行榜(2026年8月1日)上,StellaVLA以0.63总分排名第一,显著优于强基线模型π_{0.5}(0.44)和LingBot-VLA(0.22),并在LIBERO上达到98.8%平均成功率,LIBERO-Plus上达85.1%。真实机器人测试表明,StellaVLA可利用人类/机器人示范及人到机器人的(XR)示范作为上下文结构化示范,有效适应分布外任务。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often collapses out of distribution (OOD), when the scene, viewpoint, or object differs from training. Adapting to each new situation typically requires collecting more data and fine-tuning. We present StellaVLA, a framework that instead adapts at test time by conditioning on a single retrieved demonstration. The key idea is to move beyond imitating what an expert did and instead convey why: an automated offline pipeline converts each raw trajectory into a structured demonstration, e.g., a task plan, sub-goal descriptions, and verbalized 3D motion, at zero human-annotation cost. Provided as in-context guidance, this structured demonstration lets the policy reason about the task rather than mimic a pixel trajectory, which also makes it transferable across embodiments (real-robot, human-hand, or XR demonstrations). A parallel dual-training design internalizes this reasoning during training through a joint action-and-language objective, while inference uses the action expert alone, preserving real-time, high-frequency control with no added latency. On the VLA-Arena leaderboard(Aug 1, 2026), StellaVLA ranks first with an overall score of 0.63, versus 0.44 and 0.22 for the strong prior models ($π_{0.5}$ and LingBot-VLA), and it further leads on LIBERO with 98.8% average success rate and LIBERO-Plus with 85.1% success rate. Our real-robot benchmark demonstrates that StellaVLA can use both human/robot demos and human-to-robot (XR) demos as in-context structured demonstration to help VLA model adapt to OOD tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。