用视频扩散模型生成动态交互数据,训练3D物体的动态使用能力模型。
DAViD: Modeling Dynamic Affordance of 3D Objects Using Pre-trained Video Diffusion Models
- 基于预训练视频扩散模型生成2D交互视频,再升维为4D动态样本。
- 通过低秩适配微调人体运动模型,仅用少量数据学会新交互动作。
- 能融合多个交互概念生成新颖动作,适合动作生成与人机交互研究者。
理解人类如何与物体互动对AI有效辅助或模仿人类行为至关重要。现有研究多关注静态人-物交互(如接触与空间关系),而动态交互模式(涉及人与物体随时间变化的运动)仍较少被探索。本文提出DAViD框架,用于学习跨多种目标物体类别的动态使用能力。针对4D人-物交互数据集稀缺问题,方法从合成生成的4D人-物交互样本中学习。具体流程为:首先利用预训练视频扩散模型从给定3D目标物体生成2D人-物交互视频,再将其提升至3D以生成4D人-物交互样本。基于这些合成样本,训练我们的生成式4D人-物交互模型DAViD,包含两个核心组件:(1) 带有低秩适配(LoRA)模块的人体运动扩散模型(MDM),用于在有限人-物运动样本下微调预训练MDM以学习人-物交互运动概念;(2) 由生成的人体交互动作条件控制的4D物体姿态运动扩散模型。有趣的是,DAViD可将新学的交互动作概念与预训练人体动作结合,生成新颖人-物交互动作,即使同时融合多个交互概念,体现了我们基于LoRA的管道在整合动态交互概念上的优势。大量实验表明,DAViD在合成人-物交互动作方面优于基线方法。
原文摘要 · Abstract (English)
Modeling how humans interact with objects is crucial for AI to effectively assist or mimic human behaviors. Existing studies for learning such ability primarily focus on static human-object interaction (HOI) patterns, such as contact and spatial relationships, while dynamic HOI patterns, capturing the movement of humans and objects over time, remain relatively underexplored. In this paper, we present a novel framework for learning Dynamic Affordance across various target object categories. To address the scarcity of 4D HOI datasets, our method learns the 3D dynamic affordance from synthetically generated 4D HOI samples. Specifically, we propose a pipeline that first generates 2D HOI videos from a given 3D target object using a pre-trained video diffusion model, then lifts them into 3D to generate 4D HOI samples. Leveraging these synthesized 4D HOI samples, we train DAViD, our generative 4D human-object interaction model, which is composed of two key components: (1) a human motion diffusion model (MDM) with Low-Rank Adaptation (LoRA) module to fine-tune a pre-trained MDM to learn the HOI motion concepts from limited HOI motion samples, (2) a motion diffusion model for 4D object poses conditioned by produced human interaction motions. Interestingly, DAViD can integrate newly learned HOI motion concepts with pre-trained human motions to create novel HOI motions, even for multiple HOI motion concepts, demonstrating the advantage of our pipeline with LoRA in integrating dynamic HOI concepts. Through extensive experiments, we demonstrate that DAViD outperforms baselines in synthesizing HOI motion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。