用文本和少量样例让3D物体部件动起来,支持个性化动作生成。
Articulate That Object Part (ATOP): 3D Part Articulation via Text and Motion Personalization
- 基于文本提示与少量运动样本,微调扩散模型生成目标部件动作。
- 在PartNet-Mobility和ACD数据集上实现更高精度的3D动作生成。
- 适合需要快速生成特定物体部件动作的设计师或开发者。
我们提出ATOP(Articulate That Object Part),一种基于运动个性化的少样本方法,可根据文本提示对静态3D物体的特定部件进行动作生成。由于带有运动属性标注的数据集稀缺,现有方法在此任务上泛化能力不足。本文利用文本输入激活现代扩散模型,生成符合目标物体类别和部件的合理动作样本;同时,输入的3D物体作为“图像提示”,将生成动作个性化到具体对象。方法首先通过少量参考运动帧对预训练扩散模型进行微调,注入部件动作感知能力,学习与目标部件关联的独特动作标识符。微调后的模型可从多视角生成合理动作。最后,通过可微渲染将个性化动作映射至3D空间,使用得分蒸馏采样损失优化部件关节参数。在PartNet-Mobility和ACD数据集上的实验表明,该方法在少样本设置下生成的动作更真实、准确,显著提升3D动作预测的泛化能力。
原文摘要 · Abstract (English)
We present ATOP (Articulate That Object Part), a novel few-shot method based on motion personalization to articulate a static 3D object with respect to a part and its motion as prescribed in a text prompt. Given the scarcity of available datasets with motion attribute annotations, existing methods struggle to generalize well in this task. In our work, the text input allows us to tap into the power of modern-day diffusion models to generate plausible motion samples for the right object category and part. In turn, the input 3D object provides ``image prompting'' to personalize the generated motion to the very input object. Our method starts with a few-shot finetuning to inject articulation awareness to current diffusion models to learn a unique motion identifier associated with the target object part. Our finetuning is applied to a pre-trained diffusion model for controllable multi-view motion generation, trained with a small collection of reference motion frames demonstrating appropriate part motion. The resulting motion model can then be employed to realize plausible motion of the input 3D object from multiple views. At last, we transfer the personalized motion to the 3D space of the object via differentiable rendering to optimize part articulation parameters by a score distillation sampling loss. Experiments on PartNet-Mobility and ACD datasets demonstrate that our method can generate realistic motion samples with higher accuracy, leading to more generalizable 3D motion predictions compared to prior approaches in the few-shot setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。