arXiv:2411.17532cs.CV2024-11被引 4

用频域与文本状态空间模型提升文本生成动作的精度与一致性

FTMoMamba: Motion Generation with Frequency and Text State Space Models

  • 分离运动信号的高低频成分,分别建模静态姿态与精细动作
  • 在HumanML3D上实现0.181的最低FID,显著优于基线模型
  • 适合需要高精度动作生成的影视动画与虚拟人应用

扩散模型在人体动作生成中表现优异,但现有方法通常忽略潜在空间中的频域信息(如低频对应静态姿态,高频对应精细动作)。同时,文本与动作之间存在语义差异,导致生成动作与描述不符。本文提出基于扩散的FTMoMamba框架,包含频率状态空间模型(FreqSSM)和文本状态空间模型(TextSSM)。FreqSSM将序列分解为低频与高频分量,分别指导静态姿态(如坐、躺)与精细动作(如移动、踉跄)的生成;TextSSM在句子级别编码文本特征,对齐语义与时序特征。大量实验表明,FTMoMamba在文本到动作生成任务中表现卓越,尤其在HumanML3D数据集上取得0.181的最低FID,显著优于基线模型(如MLD的0.421)。

原文摘要 · Abstract (English)

Diffusion models achieve impressive performance in human motion generation. However, current approaches typically ignore the significance of frequency-domain information in capturing fine-grained motions within the latent space (e.g., low frequencies correlate with static poses, and high frequencies align with fine-grained motions). Additionally, there is a semantic discrepancy between text and motion, leading to inconsistency between the generated motions and the text descriptions. In this work, we propose a novel diffusion-based FTMoMamba framework equipped with a Frequency State Space Model (FreqSSM) and a Text State Space Model (TextSSM). Specifically, to learn fine-grained representation, FreqSSM decomposes sequences into low-frequency and high-frequency components, guiding the generation of static pose (e.g., sits, lay) and fine-grained motions (e.g., transition, stumble), respectively. To ensure the consistency between text and motion, TextSSM encodes text features at the sentence level, aligning textual semantics with sequential features. Extensive experiments show that FTMoMamba achieves superior performance on the text-to-motion generation task, especially gaining the lowest FID of 0.181 (rather lower than 0.421 of MLD) on the HumanML3D dataset.

动作生成扩散模型状态空间文本控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。