提出MoRAE,让文本生成动作更自然流畅。
MoRAE: Flow-Friendly Self-Supervised Latents for Text-to-Motion Generation

- 用压缩瓶颈保留关键运动语义,改善潜空间稳定性。
- 训练时对齐解码器敏感方向,减少解码后误差放大。
- 适合需要高精度动作生成的研究者和开发者。
文本到动作生成需保证动作语义正确、时间连贯且物理合理。现有方法通常先将运动数据投影至结构化语义空间,再在该空间中训练生成模型。此类范式在图像生成中通过表示自编码器(RAE)取得成功,其中冻结的自监督编码器为扩散或流模型提供语义特征。然而,直接将此范式应用于运动空间(以Motion-JEPA作为冻结编码器)却表现失败。我们从几何角度分析失败原因,发现两个运动特有瓶颈:(1) JEPA特征空间谱条件不良,导致从高斯分布到数据的传输不稳定;(2) 即便谱条件良好,流残差仍易沿解码器敏感方向对齐,使微小潜空间误差在解码后放大为显著动作伪影。基于此,我们提出MoRAE。MoRAE分别解决上述两个瓶颈:首先通过紧凑瓶颈压缩并去除冗余方向,使潜空间谱进入可传输稳定区间;其次通过运动耦合训练,使保留的潜空间几何与解码器对齐,降低解码后典型流误差的代价。在此流友好潜空间上,标准非自回归流匹配DiT达到当前最佳性能。
原文摘要 · Abstract (English)
Text-to-motion generation must produce motions that are semantically correct, temporally coherent, and physically plausible. A natural approach is to first project motion data into a structured semantic space and then train a generative model within that space. Such a paradigm has been highly successful in image generation through Representation Autoencoders (RAEs), where a frozen self-supervised encoder provides semantic features for diffusion or flow models to learn from. However, direct transfer of such a paradigm to motion space using Motion-JEPA as the frozen encoder fails dramatically. We diagnose this failure geometrically and identify two motion-specific bottlenecks: (1) the JEPA feature space is spectrally ill-conditioned, making the Gaussian-to-data transport unstable; and (2) even with a well-conditioned spectrum, flow residuals tend to align with decoder-sensitive directions, where small latent errors are amplified into large motion artifacts after decoding. Based on these insights, we propose MoRAE. MoRAE addresses the two bottlenecks separately. A compact bottleneck distills the structured JEPA representation while removing weak and redundant directions, bringing the latent spectrum into a transport-stable regime. Motion-coupled training then aligns the retained latent geometry with the decoder, making characteristic flow errors less costly after decoding. With this flow-friendly latent, a standard non-autoregressive Flow-Matching DiT achieves state-of-the-art performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。