用人体姿态预测第一人称视频,模拟动作如何改变环境。
Whole-Body Conditioned Egocentric Video Prediction
- 以人体关节点姿态为条件,用扩散模型预测第一人称视频。
- 在Nymeria数据集上训练,实现对真实场景中身体动作的连续视频生成。
- 设计分层评估框架,可量化分析模型的具身预测与控制能力。
我们训练模型从人体动作(PEVA)预测第一人称视频,输入包括过去视频和由相对3D身体姿态表示的动作。通过基于身体关节层次结构的运动学姿态轨迹作为条件,模型学习如何从第一人称视角模拟物理动作对环境的影响。我们在一个大规模的真实世界第一人称视频与身体姿态捕捉数据集Nymeria上,训练了一个自回归条件扩散变压器。此外,我们设计了分层评估协议,包含逐渐增加难度的任务,以全面分析模型的具身预测与控制能力。本工作是首次尝试从人类视角建模复杂现实环境及具身智能体行为的视频预测方法。
原文摘要 · Abstract (English)
We train models to Predict Ego-centric Video from human Actions (PEVA), given the past video and an action represented by the relative 3D body pose. By conditioning on kinematic pose trajectories, structured by the joint hierarchy of the body, our model learns to simulate how physical human actions shape the environment from a first-person point of view. We train an auto-regressive conditional diffusion transformer on Nymeria, a large-scale dataset of real-world egocentric video and body pose capture. We further design a hierarchical evaluation protocol with increasingly challenging tasks, enabling a comprehensive analysis of the model's embodied prediction and control abilities. Our work represents an initial attempt to tackle the challenges of modeling complex real-world environments and embodied agent behaviors with video prediction from the perspective of a human.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。