用一个模型同时生成动作和文本,效果更好还省成本。
Unlocking Pretrained LLMs for Motion-Related Multimodal Generation: A Fine-Tuning Approach to Unify Diffusion and Next-Token Prediction
- 微调预训练大模型,统一扩散生成与文本预测能力。
- 文本到动作任务上FID提升38%,七项指标准确率提高16.61%。
- 首个融合扩散与大模型生成的动作多模态框架,适合动作合成研究者。
本文提出一种统一框架MoMug,利用单一预训练大语言模型实现与动作相关的多模态生成。MoMug通过微调将基于扩散的连续动作生成与模型固有的自回归离散文本预测能力相结合,在同一模型架构中实现连续动作输出与离散文本词预测的无缝切换,有效融合了扩散模型与大语言模型的优势。实验表明,相比最新大模型基线,MoMug在文本到动作任务上将FID降低38%,七项指标平均准确率提升16.61%,八项指标平均准确率提升8.44%。据我们所知,这是首个在单模型中整合扩散与大模型生成用于动作相关多模态任务的方法,且训练成本低,为未来高质量、低成本动作合成奠定了基础。
原文摘要 · Abstract (English)
In this paper, we propose a unified framework that leverages a single pretrained LLM for Motion-related Multimodal Generation, referred to as MoMug. MoMug integrates diffusion-based continuous motion generation with the model's inherent autoregressive discrete text prediction capabilities by fine-tuning a pretrained LLM. This enables seamless switching between continuous motion output and discrete text token prediction within a single model architecture, effectively combining the strengths of both diffusion- and LLM-based approaches. Experimental results show that, compared to the most recent LLM-based baseline, MoMug improves FID by 38% and mean accuracy across seven metrics by 16.61% on the text-to-motion task. Additionally, it improves mean accuracy across eight metrics by 8.44% on the text-to-motion task. To the best of our knowledge, this is the first approach to integrate diffusion- and LLM-based generation within a single model for motion-related multimodal tasks while maintaining low training costs. This establishes a foundation for future advancements in motion-related generation, paving the way for high-quality yet cost-efficient motion synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。