机器人学习不能只靠示范数据,还需从视频、动作中提取有用信息。
Robots Need More than VLA and World Models

- 构建接口将非结构化行为数据转化为机器人可用的监督信号
- 提出四类关键接口:自动标注、身体适配、物理推理和奖励推断
- 适合关注机器人通用智能与多源学习的研究者
通用机器人智能常被看作策略规模问题:收集更多机器人示范,训练更大视觉-语言-动作(VLA)模型,期望实现更广泛泛化。本文认为这一框架不完整。核心瓶颈不仅在于策略学习,更在于缺乏将世界中丰富的非结构化行为数据转化为具身机器人监督的机制。人类动作、互联网视频、仿真轨迹及交互示范包含任务、目标、接触、失败和物理约束等丰富信息,但这些数据因缺乏特定于机器人的动作标签、任务语义和奖励结构,难以被机器人策略直接使用。本文识别出下一代机器人所需的四大缺失组件:用于自动标注非结构化行为的数据接口,用于将人类动作迁移到机器人动作的具身接口,用于物理基础3D推理的世界模型接口,以及从视频和语言中推断任务进展与成功的奖励接口。我们综述了机器人基础模型、跨具身数据集、从视频学习、世界模型和奖励建模的最新进展,并提出了一个研究议程,旨在构建不仅能从机器人示范学习,还能从更广泛的物理世界中学习的机器人系统。
原文摘要 · Abstract (English)
Generalist robot intelligence is often framed as a policy-scaling problem: collect more robot demonstrations, train larger Vision-Language-Action (VLA) models, and expect broader generalisation. In this position paper, we argue that this framing is incomplete. The central bottleneck is not only policy learning, but the absence of mechanisms that convert the world's abundant unstructured behavioural data into grounded robot supervision. Human motion, internet video, simulation rollouts, and interactive demonstrations contain rich information about tasks, goals, contacts, failures, and physical constraints, yet most of this information is not directly usable by robot policies because it lacks embodiment-specific action labels, task semantics, and reward structure. We identify four missing components for the next generation of robotics: data interfaces for autolabelling unstructured behaviour, embodiment interfaces for retargeting human motion to robot actions, world-model interfaces for physics-grounded 3D reasoning, and reward interfaces for inferring task progress and success from video and language. We survey recent progress in robot foundation models, cross-embodiment datasets, learning from video, world models, and reward modelling, and propose a research agenda for building robotics systems that can learn not only from robot demonstrations, but from the broader physical world.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。