用3D高斯与语言模型生成更真实的4D人体动作。
Object-Aware 4D Human Motion Generation
- 基于3D高斯表示和运动扩散模型,结合大模型语义信息优化动作。
- 零样本泛化,无需微调即可生成符合物体空间约束的动作。
- 适合需要真实物理交互的虚拟人像、动画生成场景。
视频扩散模型虽能生成高质量视频,但仍存在不自然形变、语义错误和物理矛盾等问题,主要源于缺乏3D物理先验。为此,我们提出基于3D高斯表示和运动扩散先验的对象感知4D人体动作生成框架。在预生成3D人体和物体的基础上,所提方法Motion Score Distilled Interaction(MSDI)利用大语言模型(LLMs)的空间与提示语义信息,通过提出的运动扩散得分蒸馏采样(MSDS),将预训练运动扩散模型的得分梯度蒸馏用于优化人体动作,使其在保持语义一致性的同时尊重物体与空间约束。与需在有限交互数据集上联合训练的方法不同,本方法为零样本设计,避免重训练,可泛化至分布外的物体相关动作生成。实验表明,该框架生成的人体动作自然且符合物理规律,充分考虑3D空间上下文,为真实4D生成提供了可扩展解决方案。
原文摘要 · Abstract (English)
Recent advances in video diffusion models have enabled the generation of high-quality videos. However, these videos still suffer from unrealistic deformations, semantic violations, and physical inconsistencies that are largely rooted in the absence of 3D physical priors. To address these challenges, we propose an object-aware 4D human motion generation framework grounded in 3D Gaussian representations and motion diffusion priors. With pre-generated 3D humans and objects, our method, Motion Score Distilled Interaction (MSDI), employs the spatial and prompt semantic information in large language models (LLMs) and motion priors through the proposed Motion Diffusion Score Distillation Sampling (MSDS). The combination of MSDS and LLMs enables our spatial-aware motion optimization, which distills score gradients from pre-trained motion diffusion models, to refine human motion while respecting object and semantic constraints. Unlike prior methods requiring joint training on limited interaction datasets, our zero-shot approach avoids retraining and generalizes to out-of-distribution object aware human motions. Experiments demonstrate that our framework produces natural and physically plausible human motions that respect 3D spatial context, offering a scalable solution for realistic 4D generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。