让文本生成动作更自然,按阶段动态调整语义细节。
ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model
- 分阶段处理:前期重结构、后期重细节,模仿生物发育规律
- 在StableMoFusion上实现最优语义对齐效果,提升生成质量
- 适合做动作生成的开发者,尤其关注细节与文本匹配的人
尽管扩散模型推动了文本到动作生成的发展,但其静态语义条件忽略了时间频率需求:早期去噪需要结构语义构建动作基础,后期则需局部细节对齐文本。这一矛盾类似生物形态发生过程中不同发育阶段所需的不同基因程序。受表观遗传调控启发,我们提出ANT(自适应神经时序感知架构),通过两个核心模块实现语义粒度的动态调控:(i) 语义时序自适应模块(STA):基于频谱分析自动划分去噪过程为低频结构规划与高频细化;(ii) 动态无分类器引导调度(DCFG):自适应调整条件与无条件比例,在保持保真度的同时提升效率。大量实验表明,ANT可适配多种基线模型,显著提升性能,并在StableMoFusion上达到当前最佳语义对齐效果。
原文摘要 · Abstract (English)
While diffusion models advance text-to-motion generation, their static semantic conditioning ignores temporal-frequency demands: early denoising requires structural semantics for motion foundations while later stages need localized details for text alignment. This mismatch mirrors biological morphogenesis where developmental phases demand distinct genetic programs. Inspired by epigenetic regulation governing morphological specialization, we propose **(ANT)**, an **A**daptive **N**eural **T**emporal-Aware architecture. ANT orchestrates semantic granularity through: **(i) Semantic Temporally Adaptive (STA) Module:** Automatically partitions denoising into low-frequency structural planning and high-frequency refinement via spectral analysis. **(ii) Dynamic Classifier-Free Guidance scheduling (DCFG):** Adaptively adjusts conditional to unconditional ratio enhancing efficiency while maintaining fidelity. Extensive experiments show that ANT can be applied to various baselines, significantly improving model performance, and achieving state-of-the-art semantic alignment on StableMoFusion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。