首个可同时生成动作与文本的扩散模型,支持双向互动生成。
PackDiT: Joint Human Motion and Text Generation via Mutual Prompting
- 用互信息块融合多模态扩散Transformer,实现跨模态协同生成。
- 在HumanML3D上达成0.106的FID,文本到动作生成领先现有方法。
- 首次证明扩散模型能媲美自回归模型完成动作到文本生成任务。
人体动作生成得益于扩散模型的进展。当前多数研究聚焦于基于文本提示的动作生成(即文本到动作),但动作与文本的双向生成——包括动作到文本和联合生成——仍鲜有探索。本文提出PackDiT,首个基于扩散模型的通用生成框架,可同时处理动作生成、动作预测、文本生成、文本到动作、动作到文本及联合生成等任务。核心创新在于通过互信息模块,无缝整合不同模态的扩散Transformer(DiTs)。在HumanML3D数据集上训练,PackDiT在文本到动作生成上取得0.106的FID,优于当前最优水平,并在动作预测与补全任务中表现优异。实验进一步表明,扩散模型在动作到文本生成上性能可媲美自回归模型。
原文摘要 · Abstract (English)
Human motion generation has advanced markedly with the advent of diffusion models. Most recent studies have concentrated on generating motion sequences based on text prompts, commonly referred to as text-to-motion generation. However, the bidirectional generation of motion and text, enabling tasks such as motion-to-text alongside text-to-motion, has been largely unexplored. This capability is essential for aligning diverse modalities and supports unconditional generation. In this paper, we introduce PackDiT, the first diffusion-based generative model capable of performing various tasks simultaneously, including motion generation, motion prediction, text generation, text-to-motion, motion-to-text, and joint motion-text generation. Our core innovation leverages mutual blocks to integrate multiple diffusion transformers (DiTs) across different modalities seamlessly. We train PackDiT on the HumanML3D dataset, achieving state-of-the-art text-to-motion performance with an FID score of 0.106, along with superior results in motion prediction and in-between tasks. Our experiments further demonstrate that diffusion models are effective for motion-to-text generation, achieving performance comparable to that of autoregressive models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。