arXiv:2506.15625cs.CV2025-06被引 18

用噪声优化生成真实自然的人物物体交互动作

HOIDiNi: Human-Object Interaction through Diffusion Noise Optimization

  • 在预训练扩散模型的噪声空间中直接优化,实现精准接触与自然动作统一
  • 在GRAB数据集上接触准确率和物理合理性均优于已有方法
  • 适合需要可控复杂交互生成的研究者或应用开发者

我们提出HOIDiNi,一种基于文本驱动的扩散框架,用于合成逼真且合理的物体交互(HOI)。HOI生成极具挑战性,需同时满足严格的接触精度与多样化的运动模式。现有方法在真实感与物理正确性之间权衡,而HOIDiNi通过扩散噪声优化(DNO)直接在预训练扩散模型的噪声空间中优化,兼顾两者。关键洞察在于将问题分解为两个阶段:以物体为中心的阶段决定手与物体的接触位置,以人体为中心的阶段细化全身动作以实现该蓝图。这一结构化方法在不牺牲动作自然性的前提下实现了精确的手物接触。在GRAB数据集上的定量、定性和主观评估表明,HOIDiNi在接触准确率、物理有效性及整体质量方面显著优于先前方法与基线。结果证明其能仅凭文本提示生成复杂、可控制的交互动作,包括抓取、放置和全身协调。

原文摘要 · Abstract (English)

We present HOIDiNi, a text-driven diffusion framework for synthesizing realistic and plausible human-object interaction (HOI). HOI generation is extremely challenging since it induces strict contact accuracies alongside a diverse motion manifold. While current literature trades off between realism and physical correctness, HOIDiNi optimizes directly in the noise space of a pretrained diffusion model using Diffusion Noise Optimization (DNO), achieving both. This is made feasible thanks to our observation that the problem can be separated into two phases: an object-centric phase, primarily making discrete choices of hand-object contact locations, and a human-centric phase that refines the full-body motion to realize this blueprint. This structured approach allows for precise hand-object contact without compromising motion naturalness. Quantitative, qualitative, and subjective evaluations on the GRAB dataset alone clearly indicate HOIDiNi outperforms prior works and baselines in contact accuracy, physical validity, and overall quality. Our results demonstrate the ability to generate complex, controllable interactions, including grasping, placing, and full-body coordination, driven solely by textual prompts. https://hoidini.github.io.

人物交互扩散模型动作生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。