无需训练即可生成多样且物理真实的3D人物交互动作
Zero-Shot Human-Object Interaction Synthesis with Multimodal Priors
- 利用预训练多模态模型提取交互知识,从文本生成2D图像序列
- 通过姿态估计与6-DoF物体定位,实现跨类别物体的3D姿态重建
- 适合需要快速生成新交互场景的研究者和开发者
人-物体交互(HOI)合成在虚拟现实、机器人等领域具有重要意义。然而,由于3D HOI数据复杂且成本高,现有方法受限于训练数据中物体类型和交互模式的狭窄多样性。本文提出一种无需端到端训练的零样本HOI合成框架,核心思想是利用预训练多模态模型中的广泛HOI知识。给定文本描述后,系统首先使用图像或视频生成模型生成时序一致的2D HOI图像序列,再将其提升为3D HOI的姿态里程碑。我们采用预训练人体姿态估计模型获取人体姿态,并引入通用的类别级6-DoF物体位姿估计方法,从2D HOI图像中恢复物体位姿。该方法对由文生3D模型或在线检索获得的多种物体模板均具适应性。进一步通过基于物理的跟踪优化3D HOI运动学里程碑,提升身体动作与物体姿态的合理性,生成更符合物理规律的交互结果。实验表明,本方法能生成开放词汇的、具备物理真实性和语义多样性的3D HOI。
原文摘要 · Abstract (English)
Human-object interaction (HOI) synthesis is important for various applications, ranging from virtual reality to robotics. However, acquiring 3D HOI data is challenging due to its complexity and high cost, limiting existing methods to the narrow diversity of object types and interaction patterns in training datasets. This paper proposes a novel zero-shot HOI synthesis framework without relying on end-to-end training on currently limited 3D HOI datasets. The core idea of our method lies in leveraging extensive HOI knowledge from pre-trained Multimodal Models. Given a text description, our system first obtains temporally consistent 2D HOI image sequences using image or video generation models, which are then uplifted to 3D HOI milestones of human and object poses. We employ pre-trained human pose estimation models to extract human poses and introduce a generalizable category-level 6-DoF estimation method to obtain the object poses from 2D HOI images. Our estimation method is adaptive to various object templates obtained from text-to-3D models or online retrieval. A physics-based tracking of the 3D HOI kinematic milestone is further applied to refine both body motions and object poses, yielding more physically plausible HOI generation results. The experimental results demonstrate that our method is capable of generating open-vocabulary HOIs with physical realism and semantic diversity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。