无需真人动作数据,直接生成人与环境的自然互动视频。
ZeroHSI: Zero-Shot 4D Human-Scene Interaction by Video Generation
- 从视频生成模型中提取人类交互知识,结合可微渲染重建4D交互
- 在静态和动态环境中都能生成真实自然的人体动作,无需真实动作数据
- 适合虚拟现实、机器人模拟等需要零样本泛化的场景
人体-场景交互(HSI)生成在具身智能、虚拟现实和机器人领域至关重要。然而,现有方法依赖于配对的3D场景与真人动作捕捉数据进行训练,无法处理未见环境(如真实场景或重建场景)。本文提出ZeroHSI,一种零样本4D人体-场景交互合成方法,无需任何动作捕捉数据训练。核心思路是利用先进视频生成模型中蕴含的自然人类动作与交互先验,并通过可微渲染重建交互过程。ZeroHSI可在静态场景及含动态物体的环境中生成逼真人体运动,且不依赖任何真实动作数据。我们在一个包含多种室内外场景及不同交互提示的精选数据集上评估,结果表明其能生成多样且符合语境的人体-场景交互。
原文摘要 · Abstract (English)
Human-scene interaction (HSI) generation is crucial for applications in embodied AI, virtual reality, and robotics. Yet, existing methods cannot synthesize interactions in unseen environments such as in-the-wild scenes or reconstructed scenes, as they rely on paired 3D scenes and captured human motion data for training, which are unavailable for unseen environments. We present ZeroHSI, a novel approach that enables zero-shot 4D human-scene interaction synthesis, eliminating the need for training on any MoCap data. Our key insight is to distill human-scene interactions from state-of-the-art video generation models, which have been trained on vast amounts of natural human movements and interactions, and use differentiable rendering to reconstruct human-scene interactions. ZeroHSI can synthesize realistic human motions in both static scenes and environments with dynamic objects, without requiring any ground-truth motion data. We evaluate ZeroHSI on a curated dataset of different types of various indoor and outdoor scenes with different interaction prompts, demonstrating its ability to generate diverse and contextually appropriate human-scene interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。