arXiv:2608.26103cs.ROcs.CV2026-08

用人类操作视频指导机器人完成未见过的任务,实现零样本泛化。

Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization

论文配图:Zero-WAM: In-Context World-Action Modeling from Human Videos for Open-Ended Task Generalization
图 1 · 摘自论文原文
  • 通过人类操作视频作为上下文提示,引导机器人执行新任务
  • 在仿真中平均成功率达47.0%,比最强基线提升29.5个百分点
  • 适用于多物体、长序列和精细插入等复杂现实任务

零样本跨任务泛化是机器人学习的核心挑战。大型语言模型可通过上下文指定新任务而无需参数更新,这种上下文学习(ICL)将泛化转化为任务描述问题。本文将其引入机器人操作,提出以人类视频作为自然的任务描述:相比语言,视频能提供丰富的任务演化视觉线索。我们提出Zero-WAM,一种因果视频-动作模型,通过遵循上下文中的人类视频指导执行未见任务。为解决任务丰富的人-机配对数据稀缺问题,构建自动化流水线,将任务采样的机器人轨迹转换为语义匹配的人类视频,生成HumanGen数据集,包含74.2K组人-机ICL对,覆盖8.6K个任务。训练时引入上下文未来片段预测(IFP)目标,抑制从已见任务中习得的捷径,强制策略从视频提示中提取任务信息。在RoboTwin 2.0仿真中,对7个未见任务,Zero-WAM平均成功率达47.0%,绝对提升29.5个百分点;真实世界测试中,可成功泛化至多物体场景、长时序操作及精细插入等新配置。

原文摘要 · Abstract (English)

Zero-shot cross-task generalization, where a policy must execute manipulation tasks never seen during training, remains a central challenge in robot learning. In large language models, a novel task can be performed simply by specifying it in the context, without any parameter update. This form of in-context learning (ICL) turns generalization into a problem of task specification. To achieve cross-task generalization, we bring this paradigm to robotic manipulation, and argue that the natural task specification for manipulation is a human video: unlike language, it provides rich visual cues about the intended task evolution. We present Zero-WAM, a causal video-action model that executes unseen tasks by following in-context human video guidance. To address the scarcity of task-rich paired human-robot data, we propose an automatic pipeline that converts task-sampled robot trajectories into semantically matched human videos, yielding HumanGen, a dataset of 74.2K human-robot ICL pairs across 8.6K tasks. For model training, we further introduce an in-context future chunk prediction (IFP) objective that suppresses shortcuts learned from seen tasks and forces the policy to draw task information from the video prompt. On seven unseen tasks in RoboTwin 2.0 simulation, Zero-WAM achieves a 47.0% average success rate, an absolute improvement of 29.5 percentage points over the strongest video-action baseline. In real-world evaluations, it follows human video guidance to generalize to unseen task configurations involving multi-object scenes, long-horizon manipulation, and fine-grained insertion.

机器人学习视频理解零样本泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。