让机器人从人类视频中学习操作,解决硬件与动作不匹配难题。
From Passive Video to Editable Experience: Physically Grounded Experience Synthesis for Embodied Intelligence

- 用任务图+功能约束图构建中间表示,实现人到机器的动作转换。
- 在多个数据集上验证,生成动作可执行且状态一致,跨机器人通用。
- 只需少量资源,即可将视频转化为机器人可学的数据,适合工业应用。
embodied AI 的核心瓶颈并非模型结构,而是数据。尽管网络上有数十亿条人类操作视频,但因人体形态与机器人硬件间的“具身鸿沟”,机器人无法直接学习。我们提出 Pegasus,一种低资源框架,通过结构化知识迁移弥合这一差距。Pegasus 不依赖原始视频提示,而是将人类视频中的任务图,经由功能图与约束图转换为机器人规划图,用于生成机器人条件视频。其分层功能潜在空间建模了物体状态、功能与任务间的关系,实现对物体身份的泛化。闭环物理验证器利用运动学可行性、碰撞约束和关节限位过滤无效生成。我们在 GTEA Gaze+ 和 EPIC-KITCHENS-100 等多个人类视角操作基准上评估,涵盖多种机器人本体,评测任务正确性、可执行性、状态一致性与可学习性。结果表明,Pegasus 能可靠实现跨本体翻译,证明机器人数据生成可从硬件采集问题转变为可扩展的低资源知识迁移问题。
原文摘要 · Abstract (English)
The key bottleneck in embodied AI is not model architecture but data. Although billions of human manipulation videos exist online, robots cannot directly learn from them due to the embodiment gap between human morphology and robot hardware. We introduce Pegasus, a low-resource framework that bridges this gap by translating human demonstrations into robot-learnable data through structured knowledge transfer. Instead of relying on raw video prompts, Pegasus constructs a graph-based intermediate representation: a Task Graph extracted from human videos is transformed through Affordance and Constraint Graphs into a Robot Planning Graph for robot-conditioned video generation. A hierarchical affordance latent space models the relationship between object states, affordances, and tasks, enabling generalization beyond object identities. A closed-loop physics verifier further filters invalid generations using kinematic feasibility, collision constraints, and joint limits. We evaluate Pegasus across a range of egocentric manipulation benchmarks, including GTEA Gaze+ and EPIC-KITCHENS-100, and diverse robot embodiments, assessing Task Correctness, Executability, State Consistency, and Learnability. Results demonstrate reliable cross-embodiment translation and show that robot data generation can be reframed from a hardware collection problem into a scalable, low-resource knowledge transfer problem.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。