arXiv:2502.05534cs.CV2025-02IJCV被引 52

用大模型精准解析文本,生成更符合描述的人体动作。

Fg-T2M++: LLMs-Augmented Fine-Grained Text Driven Human Motion Generation

  • 用大模型解析文本中身体部位的细节语义
  • 在双曲空间建模词语间关系,提升语义理解
  • 适合需要精确动作控制的影视与游戏应用

我们解决细粒度文本驱动人体动作生成的挑战。现有方法因缺乏对身体部位细节语义的有效解析,以及未能充分建模词语间的语言结构,导致生成动作不精确。为此,我们提出Fg-T2M++框架:(1) 基于大模型的语义解析模块,从文本中提取身体部位描述与语义;(2) 双曲文本表示模块,将句法依存图嵌入双曲空间以编码文本单元间的关联信息;(3) 多模态融合模块,分层融合文本与动作特征。在HumanML3D和KIT-ML数据集上的大量实验表明,Fg-T2M++优于当前最优方法,验证了其生成符合完整文本语义的动作的能力。

原文摘要 · Abstract (English)

We address the challenging problem of fine-grained text-driven human motion generation. Existing works generate imprecise motions that fail to accurately capture relationships specified in text due to: (1) lack of effective text parsing for detailed semantic cues regarding body parts, (2) not fully modeling linguistic structures between words to comprehend text comprehensively. To tackle these limitations, we propose a novel fine-grained framework Fg-T2M++ that consists of: (1) an LLMs semantic parsing module to extract body part descriptions and semantics from text, (2) a hyperbolic text representation module to encode relational information between text units by embedding the syntactic dependency graph into hyperbolic space, and (3) a multi-modal fusion module to hierarchically fuse text and motion features. Extensive experiments on HumanML3D and KIT-ML datasets demonstrate that Fg-T2M++ outperforms SOTA methods, validating its ability to accurately generate motions adhering to comprehensive text semantics.

动作生成大模型文本驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。