GEM可精准控制动作、物体和视角,生成长时序多模态视频。
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control
- 基于参考帧与人体姿态等多模态信号,实现对运动与场景的精细控制
- 支持4000+小时跨域数据训练,生成长序列视频且保持时序一致性
- 适用于自动驾驶、第一人称动作建模等需高精度控制的场景
我们提出GEM,一种通用的自视点多模态世界模型,利用参考帧、稀疏特征、人体姿态和自车轨迹预测未来帧。该模型可精确控制物体动态、自车运动及人体姿态。GEM生成配对的RGB与深度输出,增强空间理解能力。通过自回归噪声调度实现稳定长时序生成。数据集包含4000+小时多模态数据,涵盖自动驾驶、第一人称人类活动和无人机飞行等场景。伪标签用于获取深度图、自车轨迹和人体姿态。我们设计了新的物体操作控制(COM)评估指标,结合全面评估框架验证可控性。实验表明,GEM在生成多样化、可控且长时间一致的场景方面表现优异。代码、模型与数据集均开源。
原文摘要 · Abstract (English)
We present GEM, a Generalizable Ego-vision Multimodal world model that predicts future frames using a reference frame, sparse features, human poses, and ego-trajectories. Hence, our model has precise control over object dynamics, ego-agent motion and human poses. GEM generates paired RGB and depth outputs for richer spatial understanding. We introduce autoregressive noise schedules to enable stable long-horizon generations. Our dataset is comprised of 4000+ hours of multimodal data across domains like autonomous driving, egocentric human activities, and drone flights. Pseudo-labels are used to get depth maps, ego-trajectories, and human poses. We use a comprehensive evaluation framework, including a new Control of Object Manipulation (COM) metric, to assess controllability. Experiments show GEM excels at generating diverse, controllable scenarios and temporal consistency over long generations. Code, models, and datasets are fully open-sourced.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。