让文字生成动作更连贯,解决不同动作间时间结构不一致问题。
Temporal Consistency-Aware Text-to-Motion Generation
- 用跨序列时间对齐机制,统一相同动作的时间结构。
- 在HumanML3D和KIT-ML上达到当前最好性能,动作更自然合理。
- 适合做动作生成、动画合成的研究者与开发者参考。
文本到动作(T2M)生成旨在从自然语言描述中合成逼真的动作序列。尽管基于离散动作表示的两阶段框架推动了该领域的发展,但它们常忽略跨序列时间一致性,即同一动作在不同实例间共享的时间结构。这导致语义错位和物理上不合理的动作。为此,我们提出TCA-T2M,一种面向时间一致性的文本到动作生成框架。该方法引入时间一致性感知的空间VQ-VAE(TCaS-VQ-VAE),实现跨序列时间对齐,并结合掩码动作Transformer进行文本条件动作生成。此外,运动学约束模块有效缓解离散化带来的伪影,确保动作物理合理性。在HumanML3D和KIT-ML基准上的实验表明,TCA-T2M取得当前最优性能,凸显时间一致性对鲁棒且连贯的T2M生成的重要性。
原文摘要 · Abstract (English)
Text-to-Motion (T2M) generation aims to synthesize realistic human motion sequences from natural language descriptions. While two-stage frameworks leveraging discrete motion representations have advanced T2M research, they often neglect cross-sequence temporal consistency, i.e., the shared temporal structures present across different instances of the same action. This leads to semantic misalignments and physically implausible motions. To address this limitation, we propose TCA-T2M, a framework for temporal consistency-aware T2M generation. Our approach introduces a temporal consistency-aware spatial VQ-VAE (TCaS-VQ-VAE) for cross-sequence temporal alignment, coupled with a masked motion transformer for text-conditioned motion generation. Additionally, a kinematic constraint block mitigates discretization artifacts to ensure physical plausibility. Experiments on HumanML3D and KIT-ML benchmarks demonstrate that TCA-T2M achieves state-of-the-art performance, highlighting the importance of temporal consistency in robust and coherent T2M generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。