arXiv:2410.13911cs.CV2024-10中稿 · WACV 2026被引 7

生成手物交互的逼真人体姿态,控制物体相对位置。

GraspDiffusion: Synthesizing Realistic Whole-body Hand-Object Interaction

  • 分步生成人体与手部姿态,联合优化为抓握动作。
  • 可生成多样且真实的全身手物交互图像。
  • 适合做虚拟场景、人机交互数据生成的研究者。

当前生成模型虽能合成高质量图像,但在生成人体用手与物体互动方面表现不佳,主要因对交互理解不足及身体复杂区域建模困难。本文提出GraspDiffusion,一种新型生成方法,可生成真实感强的人体-物体交互场景。给定3D物体,该方法通过分别利用人体和手部姿态的生成先验,优化生成联合抓握姿态,控制物体相对于人体的位置。此姿态引导图像合成,准确反映预期交互,生成多样化且逼真的交互图像。实验表明,GraspDiffusion成功解决全身体互动生成这一较少研究的问题,优于先前方法。

原文摘要 · Abstract (English)

Recent generative models can synthesize high-quality images, but they often fail to generate humans interacting with objects using their hands. This arises mostly from the model's misunderstanding of such interactions and the hardships of synthesizing intricate regions of the body. In this paper, we propose \textbf{GraspDiffusion}, a novel generative method that creates realistic scenes of human-object interaction. Given a 3D object, GraspDiffusion constructs whole-body poses with control over the object's location relative to the human body, which is achieved by separately leveraging the generative priors for body and hand poses, optimizing them into a joint grasping pose. This pose guides the image synthesis to correctly reflect the intended interaction, creating realistic and diverse human-object interaction scenes. We demonstrate that GraspDiffusion can successfully tackle the relatively uninvestigated problem of generating full-bodied human-object interactions while outperforming previous methods. Our project page is available at https://yj7082126.github.io/graspdiffusion/

手物交互生成模型姿态生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。