用大模型精准解析文本,生成更符合描述的人体动作。
Fg-T2M++: LLMs-Augmented Fine-Grained Text Driven Human Motion Generation
- 用大模型解析文本中身体部位的细节语义
- 在双曲空间建模词语间关系,提升语义理解
- 适合需要精确动作控制的影视与游戏应用
我们解决细粒度文本驱动人体动作生成的挑战。现有方法因缺乏对身体部位细节语义的有效解析,以及未能充分建模词语间的语言结构,导致生成动作不精确。为此,我们提出Fg-T2M++框架:(1) 基于大模型的语义解析模块,从文本中提取身体部位描述与语义;(2) 双曲文本表示模块,将句法依存图嵌入双曲空间以编码文本单元间的关联信息;(3) 多模态融合模块,分层融合文本与动作特征。在HumanML3D和KIT-ML数据集上的大量实验表明,Fg-T2M++优于当前最优方法,验证了其生成符合完整文本语义的动作的能力。
原文摘要 · Abstract (English)
We address the challenging problem of fine-grained text-driven human motion generation. Existing works generate imprecise motions that fail to accurately capture relationships specified in text due to: (1) lack of effective text parsing for detailed semantic cues regarding body parts, (2) not fully modeling linguistic structures between words to comprehend text comprehensively. To tackle these limitations, we propose a novel fine-grained framework Fg-T2M++ that consists of: (1) an LLMs semantic parsing module to extract body part descriptions and semantics from text, (2) a hyperbolic text representation module to encode relational information between text units by embedding the syntactic dependency graph into hyperbolic space, and (3) a multi-modal fusion module to hierarchically fuse text and motion features. Extensive experiments on HumanML3D and KIT-ML datasets demonstrate that Fg-T2M++ outperforms SOTA methods, validating its ability to accurately generate motions adhering to comprehensive text semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。