arXiv:2409.08278cs.CV2024-09ICCV被引 10

让3D人体根据文字描述真实互动任意物体,无需大量训练数据。

DreamHOI: Subject-Driven Generation of 3D Human-Object Interactions with Diffusion Priors

  • 用扩散模型生成图像梯度,优化人体动作姿态
  • 通过隐式(NeRF)与显式(骨骼驱动)结合实现精准交互生成
  • 零样本生成新交互,适合动画、游戏与虚拟仿真领域

我们提出 DreamHOI,一种用于零样本合成人-物交互(HOIs)的新方法,使3D人体模型能根据文本描述与任意物体真实互动。该任务因真实物体类别和几何形状多样,且缺乏涵盖丰富交互的数据集而复杂。为避免依赖大量数据,我们利用在数十亿图像-文本对上训练的文生图扩散模型。通过分数蒸馏采样(SDS)获取这些模型预测的图像空间梯度,优化带皮肤的人体网格的运动参数。然而,直接将图像空间梯度反向传播至复杂的运动参数无效,因其局部性。为此,我们引入人体网格的双隐式-显式表示:结合隐式神经辐射场(NeRF)与显式骨骼驱动网格变形。优化过程中动态切换表示形式,在保持NeRF生成质量的同时,逐步精炼网格动作。通过大量实验验证,方法在生成真实感人-物交互方面表现优异。

原文摘要 · Abstract (English)

We present DreamHOI, a novel method for zero-shot synthesis of human-object interactions (HOIs), enabling a 3D human model to realistically interact with any given object based on a textual description. This task is complicated by the varying categories and geometries of real-world objects and the scarcity of datasets encompassing diverse HOIs. To circumvent the need for extensive data, we leverage text-to-image diffusion models trained on billions of image-caption pairs. We optimize the articulation of a skinned human mesh using Score Distillation Sampling (SDS) gradients obtained from these models, which predict image-space edits. However, directly backpropagating image-space gradients into complex articulation parameters is ineffective due to the local nature of such gradients. To overcome this, we introduce a dual implicit-explicit representation of a skinned mesh, combining (implicit) neural radiance fields (NeRFs) with (explicit) skeleton-driven mesh articulation. During optimization, we transition between implicit and explicit forms, grounding the NeRF generation while refining the mesh articulation. We validate our approach through extensive experiments, demonstrating its effectiveness in generating realistic HOIs.

3D生成扩散模型人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。