arXiv:2606.24884cs.ROcs.AI2026-06

让视觉语言动作模型自主学习新操作技能,无需人类示范。

InSight: Self-Guided Skill Acquisition via Steerable VLAs

论文配图:InSight: Self-Guided Skill Acquisition via Steerable VLAs
图 1 · 摘自论文原文
  • 通过分解演示视频和末端执行器位置,自动识别基础动作单元。
  • 在仿真和真实场景中成功学会翻块、开抽屉等新技能,无需人工示范。
  • 适合希望实现机器人持续学习新操作的科研与工程人员。

视觉-语言-动作(VLA)模型可从示范中学习操作技能,但其能力受限于训练数据中的技能范围。本文提出InSight框架,通过在基础动作级别(如“将夹爪移至碗上”、“向上抬起”、“倒瓶”)实现对VLA的可控引导,解锁自主技能获取。该框架包含两个阶段:(1) 自动化分割流程,利用视觉语言模型(VLM)进行计划分解并结合末端执行器位姿,将示范划分为带标签的基础动作;(2) 基于VLM的数据飞轮机制,识别完成新任务所需缺失的动作,自动尝试以VLM提出的低层控制指令进行示范,并自动标注、存储和整合成功示范至VLA训练集。我们在仿真与真实世界操作任务中评估了InSight,涵盖翻转积木、关闭抽屉、清扫、扭转和倾倒等任务,均未使用目标技能的人类示范。学成后,这些基础动作可组合执行新型长时程任务,无需额外人工示范。结果表明,基础动作可控性为VLA策略的持续技能获取提供了可行基础。

原文摘要 · Abstract (English)

Vision-language-action (VLA) models can learn manipulation skills from demonstrations, but their capabilities are bounded by the skills in the training data. We present InSight, a framework that unlocks autonomous skill acquisition by rendering VLAs steerable at the primitive-action level (e.g., "move gripper to the bowl", "lift upward", "pour the bottle"). InSight consists of two primary stages: (1) an automated segmentation pipeline that partitions demonstrations into labeled primitives via VLM plan decomposition and end-effector poses to enable VLA primitive steerability, and (2) a VLM-guided data flywheel that identifies missing primitives required to accomplish a novel task, autonomously attempts demonstrations of the missing primitives with VLM-proposed low-level control, and automatically labels, stores, and integrates successful demonstrations into the VLA training set. We evaluate InSight across simulation and real-world manipulation tasks, including block flipping, drawer closing, sweeping, twisting, and pouring, without any human demonstrations of these target skills. Once learned, these primitives can be composed to execute novel, long-horizon tasks without additional human demonstrations. Our findings demonstrate that primitive steerability provides a practical foundation for continual skill acquisition in VLA policies. Project website: https://insight-vla.github.io.

机器人技能学习自监督多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。