用动作条件控制生成真实的人物-物体交互视频。
LOME: Learning Human-Object Manipulation with Action-Conditioned Egocentric World Model
- 基于视觉和动作输入,联合建模人体动作与环境上下文。
- 可精准模拟倒水等物理交互,视频时序一致性优于现有方法。
- 适合虚拟现实、机器人训练等需真实交互的场景。
人类-物体操作的学习因动作精细且接触密集而面临挑战。传统物理驱动动画需大量建模与手工设置,且难以泛化到不同物体形态或真实环境。为此,我们提出 LOME,一种以输入图像、文本提示及每帧人体动作(包括姿态与手势)为条件的自中心世界模型,可生成真实的人体-物体交互视频。LOME 在训练中联合估计空间人体动作与环境上下文,注入强而精确的动作引导。在多样自中心人体-物体交互视频上微调预训练视频生成模型后,LOME 不仅实现高动作跟随精度与对未见场景的良好泛化能力,还能真实再现手-物交互的物理后果,例如执行“倒水”动作后液体从瓶中流入杯中。大量实验表明,该视频框架在时间一致性和运动控制方面显著优于当前最先进的图像/视频条件生成方法及图像/文本到视频生成模型。LOME 为逼真的增强现实/虚拟现实体验和可扩展的机器人训练铺平道路,无需依赖模拟环境或显式3D/4D建模。
原文摘要 · Abstract (English)
Learning human-object manipulation presents significant challenges due to its fine-grained and contact-rich nature of the motions involved. Traditional physics-based animation requires extensive modeling and manual setup, and more importantly, it neither generalizes well across diverse object morphologies nor scales effectively to real-world environment. To address these limitations, we introduce LOME, an egocentric world model that can generate realistic human-object interactions as videos conditioned on an input image, a text prompt, and per-frame human actions, including both body poses and hand gestures. LOME injects strong and precise action guidance into object manipulation by jointly estimating spatial human actions and the environment contexts during training. After finetuning a pretrained video generative model on videos of diverse egocentric human-object interactions, LOME demonstrates not only high action-following accuracy and strong generalization to unseen scenarios, but also realistic physical consequences of hand-object interactions, e.g., liquid flowing from a bottle into a mug after executing a ``pouring'' action. Extensive experiments demonstrate that our video-based framework significantly outperforms state-of-the-art image based and video-based action-conditioned methods and Image/Text-to-Video (I/T2V) generative model in terms of both temporal consistency and motion control. LOME paves the way for photorealistic AR/VR experiences and scalable robotic training, without being limited to simulated environments or relying on explicit 3D/4D modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。