用多部件检索增强生成,让文字生成动作更准更泛化。
MoRAG -- Multi-Fusion Retrieval Augmented Generation for Human Motion
- 通过多部件检索融合,提升文本到动作生成的准确性。
- 在多个数据集上显著提升运动扩散模型性能,实现即插即用。
- 适合需要高精度动作生成的影视动画与虚拟人场景。
我们提出MoRAG,一种基于多部件融合的检索增强生成方法,用于文本驱动的人体动作生成。该方法通过改进的动作检索流程获取额外知识,增强运动扩散模型的表现。通过有效提示大型语言模型(LLMs),解决动作检索中的拼写错误和语义重述问题。采用多部件检索策略,提升动作检索在语言空间中的泛化能力。通过组合检索到的动作片段,生成多样化动作样本。同时,利用低层级、部位特异的运动信息,可为未见文本描述构建合理动作。实验表明,该框架可作为即插即用模块,显著提升运动扩散模型性能。代码、预训练模型及示例视频已公开:https://motion-rag.github.io/
原文摘要 · Abstract (English)
We introduce MoRAG, a novel multi-part fusion based retrieval-augmented generation strategy for text-based human motion generation. The method enhances motion diffusion models by leveraging additional knowledge obtained through an improved motion retrieval process. By effectively prompting large language models (LLMs), we address spelling errors and rephrasing issues in motion retrieval. Our approach utilizes a multi-part retrieval strategy to improve the generalizability of motion retrieval across the language space. We create diverse samples through the spatial composition of the retrieved motions. Furthermore, by utilizing low-level, part-specific motion information, we can construct motion samples for unseen text descriptions. Our experiments demonstrate that our framework can serve as a plug-and-play module, improving the performance of motion diffusion models. Code, pretrained models and sample videos are available at: https://motion-rag.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。