让人体与物体互动更真实,通过物理模拟生成动态4D场景。
PhyGenHOI: Physically-Aware 4D Generation of Dynamic Human-Object Interactions

- 用扩散模型生成人体动作,材料点法模拟物体物理响应。
- 生成动作与物体接触时符合物理规律,交互更自然。
- 适合做虚拟试穿、动画生成或人机交互研究的开发者。
我们解决生成具有物理准确性和视觉真实性的4D人-物交互问题。给定由3D高斯斑点(3DGS)表示的静态人体和目标物体,目标是根据输入文本合成动态场景,使人体以击打、踢击等动作主动与物体交互。为此,我们提出PhyGenHOI框架,将生成式人体运动与显式物理对象模拟相结合。人体建模为由运动扩散模型(MDM)驱动的语义智能体,物体则通过材料点法(MPM)进行物理仿真,均采用3D高斯作为统一且可微的表示。通过三种耦合机制监督交互:(1) 窗口吸引损失,实现生成动作在时间上与物体相遇;(2) 接触触发重模拟步骤,确保碰撞时动量传递符合物理;(3) 掩码视频-SDS目标,注入视频先验提升接触保真度。实验表明,PhyGenHOI在多种动作、人体与物体组合下生成了物理一致的4D HOI,优于基线方法。
原文摘要 · Abstract (English)
We address the task of generating physically accurate and visually faithful 4D Human-Object Interaction (HOI). Given a static 3D human and target object represented as 3D Gaussian Splats (3DGS), our goal is to synthesize dynamic scenes where the human actively engages with the object through actions, such as punching or kicking, in accordance with a given input text. To this end, we introduce PhyGenHOI, a novel framework that couples generative human motion with an explicit physical object simulation. We model the human as a semantic agent driven by a Motion Diffusion Model (MDM) and the object as a physical agent simulated via the Material Point Method (MPM), utilizing 3D Gaussians as a unified, differentiable representation. We supervise their interaction through three coupled mechanisms: (1) A Windowed Attraction Loss that temporally synchronizes generative motion to intercept the object; (2) A Contact-Driven Re-simulation step that triggers physically consistent momentum transfer upon impact; and (3) A Masked Video-SDS objective that injects video-based priors to enhance contact fidelity. Experiments show PhyGenHOI generates physically consistent 4D HOI across diverse actions, humans, and objects, outperforming baselines. Project page and videos: https://omerbenishu.github.io/PhyGenHOI/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。