统一多粒度动作理解与生成,支持细粒度动作控制
MG-MotionLLM: A Unified Framework for Motion Comprehension and Generation across Multiple Granularities
- 构建多粒度动作-文本联合建模框架,支持粗细粒度对齐
- 引入定位动作片段边界与细节描述等辅助任务,提升模型泛化能力
- 适用于动作编辑、局部控制等新场景,适合动作生成与理解研究者
近期的动作感知大语言模型在统一动作理解与生成方面展现出潜力,但现有方法主要聚焦于粗粒度动作-文本建模,即用少数词语描述整个动作序列的整体语义,难以处理特定身体部位的细粒度动作任务。为此,我们提出 MG-MotionLLM,一个面向多粒度动作理解与生成的统一模型。通过引入一套新颖的辅助训练任务,包括基于详细文本定位动作片段的时间边界以及动作细节描述生成,实现不同粒度间动作-文本建模的相互增强。大量实验表明,MG-MotionLLM 在经典的文本到动作和动作到文本任务上表现优异,并在新的细粒度动作理解与编辑任务中展现出潜力。
原文摘要 · Abstract (English)
Recent motion-aware large language models have demonstrated promising potential in unifying motion comprehension and generation. However, existing approaches primarily focus on coarse-grained motion-text modeling, where text describes the overall semantics of an entire motion sequence in just a few words. This limits their ability to handle fine-grained motion-relevant tasks, such as understanding and controlling the movements of specific body parts. To overcome this limitation, we pioneer MG-MotionLLM, a unified motion-language model for multi-granular motion comprehension and generation. We further introduce a comprehensive multi-granularity training scheme by incorporating a set of novel auxiliary tasks, such as localizing temporal boundaries of motion segments via detailed text as well as motion detailed captioning, to facilitate mutual reinforcement for motion-text modeling across various levels of granularity. Extensive experiments show that our MG-MotionLLM achieves superior performance on classical text-to-motion and motion-to-text tasks, and exhibits potential in novel fine-grained motion comprehension and editing tasks. Project page: CVI-SZU/MG-MotionLLM
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。