用混合表示统一文本到动作生成,兼顾语义与细节。
FlowCoMotion: Text-to-Motion Generation via Token-Latent Flow Modeling

- 引入符号-隐变量耦合机制,融合离散语义与连续动态
- 在HumanML3D和SnapMoGen上达到领先性能
- 适合需要高保真动作生成的研究者
文本到动作生成依赖于运动表示与语言的语义对齐。现有方法采用连续或离散的运动表示,但连续表示混淆语义与动力学,离散表示则丢失细粒度动作细节。为此,我们提出FlowCoMotion,一种从建模视角统一两种处理的新框架。具体地,通过符号-隐变量耦合捕捉语义内容与高保真运动细节:在隐变量分支中使用多视图蒸馏正则化连续隐空间;在符号分支中采用离散时间分辨率量化提取高层语义线索。通过符号-隐变量耦合网络融合两分支表示,得到运动隐变量。随后基于文本条件预测速度场,利用常微分方程求解器从简单先验积分该速度场,引导样本到达目标动作的潜在状态。大量实验表明,FlowCoMotion在HumanML3D和SnapMoGen等文本到动作基准上表现优异。
原文摘要 · Abstract (English)
Text-to-motion generation is driven by learning motion representations for semantic alignment with language. Existing methods rely on either continuous or discrete motion representations. However, continuous representations entangle semantics with dynamics, while discrete representations lose fine-grained motion details. In this context, we propose FlowCoMotion, a novel motion generation framework that unifies both treatments from a modeling perspective. Specifically, FlowCoMotion employs token-latent coupling to capture both semantic content and high-fidelity motion details. In the latent branch, we apply multi-view distillation to regularize the continuous latent space, while in the token branch we use discrete temporal resolution quantization to extract high-level semantic cues. The motion latent is then obtained by combining the representations from the two branches through a token-latent coupling network. Subsequently, a velocity field is predicted based on the textual conditions. An ODE solver integrates this velocity field from a simple prior, thereby guiding the sample to the potential state of the target motion. Extensive experiments show that FlowCoMotion achieves competitive performance on text-to-motion benchmarks, including HumanML3D and SnapMoGen.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。