让文字生成多样物体的逼真动作,支持从未见过的物种结构。
How to Move Your Dragon: Text-to-Motion Synthesis for Large-Vocabulary Objects
- 用文本描述增强动物动作数据集,支持跨物种控制。
- 通过骨架增强技术生成多样动作,保持动力学一致性。
- 动态适配任意骨骼结构,实现异构对象的动作合成。
多样物体的动作合成在3D内容创作中潜力巨大,但因两大挑战而未被充分探索:(1)缺乏涵盖广泛高质量动作与标注的综合性动作数据集;(2)现有方法难以处理来自不同物体的异构骨骼模板。为此,我们提出三项贡献:首先,我们扩展了包含70多个物种的Truebones Zoo高质量动物动作数据集,加入详细文本描述,使其适用于基于文本的动作合成;其次,引入骨架增强技术,在保持一致动力学的前提下生成多样化运动数据,使模型可适应多种骨骼配置;最后,重新设计现有的运动扩散模型,使其能够动态适配任意骨骼模板,从而实现对具有不同结构的多样化物体进行动作合成。实验表明,该方法能从文本描述中生成高保真动作,适用于多种甚至未见物体,为跨类别、跨骨骼结构的动作合成奠定了坚实基础。定性结果详见:https://t2m4lvo.github.io
原文摘要 · Abstract (English)
Motion synthesis for diverse object categories holds great potential for 3D content creation but remains underexplored due to two key challenges: (1) the lack of comprehensive motion datasets that include a wide range of high-quality motions and annotations, and (2) the absence of methods capable of handling heterogeneous skeletal templates from diverse objects. To address these challenges, we contribute the following: First, we augment the Truebones Zoo dataset, a high-quality animal motion dataset covering over 70 species, by annotating it with detailed text descriptions, making it suitable for text-based motion synthesis. Second, we introduce rig augmentation techniques that generate diverse motion data while preserving consistent dynamics, enabling models to adapt to various skeletal configurations. Finally, we redesign existing motion diffusion models to dynamically adapt to arbitrary skeletal templates, enabling motion synthesis for a diverse range of objects with varying structures. Experiments show that our method learns to generate high-fidelity motions from textual descriptions for diverse and even unseen objects, setting a strong foundation for motion synthesis across diverse object categories and skeletal templates. Qualitative results are available at: $\href{https://t2m4lvo.github.io}{https://t2m4lvo.github.io}$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。