arXiv:2411.18660cs.CV2024-11被引 8

文本驱动生成3D人体与物体交互,支持未见过的物体和动作。

OOD-HOI: Text-Driven 3D Whole-Body Human-Object Interactions Generation Beyond Training Domains

  • 双分支循环扩散模型生成初始交互姿态
  • 接触区域引导优化提升物理合理性
  • 动态适应机制增强对新场景的泛化能力

从文本描述生成逼真的3D人体-物体交互(HOIs)是虚拟现实、机器人和动画等领域的活跃研究方向。然而,由于缺乏大规模交互数据且难以保证物理合理性,尤其在分布外(OOD)场景下,高质量3D HOI生成仍具挑战。现有方法多聚焦于身体或手部,难以实现整体协调的自然交互。本文提出OOD-HOI,一种文本驱动的全身体交互生成框架,可有效泛化至新物体与新动作。该方法结合双分支循环扩散模型生成初始交互姿态,利用接触区域引导的交互精炼模块提升物理准确性,并引入动态适应机制(包含语义调整与几何变形)以增强鲁棒性。实验表明,相比现有方法,本方案在分布外场景下生成的3D交互姿态更真实、更符合物理规律。

原文摘要 · Abstract (English)

Generating realistic 3D human-object interactions (HOIs) from text descriptions is a active research topic with potential applications in virtual and augmented reality, robotics, and animation. However, creating high-quality 3D HOIs remains challenging due to the lack of large-scale interaction data and the difficulty of ensuring physical plausibility, especially in out-of-domain (OOD) scenarios. Current methods tend to focus either on the body or the hands, which limits their ability to produce cohesive and realistic interactions. In this paper, we propose OOD-HOI, a text-driven framework for generating whole-body human-object interactions that generalize well to new objects and actions. Our approach integrates a dual-branch reciprocal diffusion model to synthesize initial interaction poses, a contact-guided interaction refiner to improve physical accuracy based on predicted contact areas, and a dynamic adaptation mechanism which includes semantic adjustment and geometry deformation to improve robustness. Experimental results demonstrate that our OOD-HOI could generate more realistic and physically plausible 3D interaction pose in OOD scenarios compared to existing methods.

3D生成文本生成人体交互泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。