arXiv:2512.07582cs.RO2025-12被引 5

仅看一次人类操作视频,机器人就能学会新任务。

See Once, Then Act: Vision-Language-Action Model with Task Learning from One-Shot Video Demonstrations

  • 通过视觉-语言-动作联合建模,从单次演示视频中提取精细操作知识。
  • 在未见过的任务上提升超30%,跨机器人形态仍保持35%以上增益。
  • 基于真实人类视频生成海量训练数据,适合快速部署到新场景的机器人系统。

构建鲁棒且通用的抓取策略是机器人研究的核心目标。尽管视觉-语言-动作(VLA)模型在端到端机器人控制方面展现出潜力,但现有方法对训练分布外任务的泛化能力仍然有限。相比之下,人类只需观察一次他人操作即可掌握新技能。受此启发,我们提出ViVLA,一种能在测试时仅依赖单次专家演示视频实现高效任务学习的通用机器人操作策略。该方法联合处理专家演示视频与机器人的视觉观测,同时预测演示的动作序列与后续机器人动作,从而将专家行为中的细粒度操作知识有效提炼并无缝迁移至智能体。为提升性能,我们设计了一条可扩展的专家-智能体配对数据生成流水线,可从易获取的人类视频中合成成对轨迹,并进一步融合公开数据集中的精选样本,共生成892,911个专家-智能体样本用于训练ViVLA。实验结果表明,该方法仅需一次演示视频即可习得新技能,在未见的LIBERO任务上提升超过30%,跨机器人形态视频下仍保持35%以上收益。真实世界实验验证了从人类视频中学习的有效性,使未见任务表现提升超过38%。

原文摘要 · Abstract (English)

Developing robust and general-purpose manipulation policies represents a fundamental objective in robotics research. While Vision-Language-Action (VLA) models have demonstrated promising capabilities for end-to-end robot control, existing approaches still exhibit limited generalization to tasks beyond their training distributions. In contrast, humans possess remarkable proficiency in acquiring novel skills by simply observing others performing them once. Inspired by this capability, we propose ViVLA, a generalist robotic manipulation policy that achieves efficient task learning from a single expert demonstration video at test time. Our approach jointly processes an expert demonstration video alongside the robot's visual observations to predict both the demonstrated action sequences and subsequent robot actions, effectively distilling fine-grained manipulation knowledge from expert behavior and transferring it seamlessly to the agent. To enhance the performance of ViVLA, we develop a scalable expert-agent pair data generation pipeline capable of synthesizing paired trajectories from easily accessible human videos, further augmented by curated pairs from publicly available datasets. This pipeline produces a total of 892,911 expert-agent samples for training ViVLA. Experimental results demonstrate that our ViVLA is able to acquire novel manipulation skills from only a single expert demonstration video at test time. Our approach achieves over 30% improvement on unseen LIBERO tasks and maintains above 35% gains with cross-embodiment videos. Real-world experiments demonstrate effective learning from human videos, yielding more than 38% improvement on unseen tasks.

机器人操控少样本学习视觉-语言-动作单次示范

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。