通过频域分析提升人形机器人文本生成动作的稳定性和语义准确性
Free-T2M: Robust Text-to-Motion Generation for Humanoid Robots via Frequency-Domain
- 从频域视角重构文本到动作生成,分阶段建模全局轨迹与细节执行
- 在StableMoFusion基础上将FID从0.152降至0.060,显著提升运动质量
- 适合关注机器人动作生成、扩散模型优化的研究者和开发者
让人形机器人从自然语言指令中合成复杂且物理合理的动作,是自主机器人与人机交互的核心挑战。尽管扩散模型在文本到动作(T2M)任务中展现出潜力,但常生成语义错误或不稳定的动作,限制了其在真实机器人的应用。本文从频域角度重新审视T2M问题,发现生成过程类似分层控制:低频成分负责建立全局运动轨迹(语义规划阶段),高频成分则细化具体动作(精细执行阶段)。为此,我们提出频域增强的文本到动作框架Free-T2M,引入阶段特异性频域一致性对齐机制,并设计频域时序自适应模块,动态调节不同频段的对齐效果。该设计强化了基础语义规划的鲁棒性,提升了细节执行的精确度。大量实验表明,方法显著改善动作质量和语义正确性。尤其在StableMoFusion基线上的FID从0.152降至0.060,成为扩散架构下的新基准。研究凸显频域洞察在生成可靠动作中的关键作用,为更直观的自然语言机器人控制铺平道路。
原文摘要 · Abstract (English)
Enabling humanoid robots to synthesize complex, physically coherent motions from natural language commands is a cornerstone of autonomous robotics and human-robot interaction. While diffusion models have shown promise in this text-to-motion (T2M) task, they often generate semantically flawed or unstable motions, limiting their applicability to real-world robots. This paper reframes the T2M problem from a frequency-domain perspective, revealing that the generative process mirrors a hierarchical control paradigm. We identify two critical phases: a semantic planning stage, where low-frequency components establish the global motion trajectory, and a fine-grained execution stage, where high-frequency details refine the movement. To address the distinct challenges of each phase, we introduce Frequency enhanced text-to-motion (Free-T2M), a framework incorporating stage-specific frequency-domain consistency alignment. We design a frequency-domain temporal-adaptive module to modulate the alignment effects of different frequency bands. These designs enforce robustness in the foundational semantic plan and enhance the accuracy of detailed execution. Extensive experiments show our method dramatically improves motion quality and semantic correctness. Notably, when applied to the StableMoFusion baseline, Free-T2M reduces the FID from 0.152 to 0.060, establishing a new state-of-the-art within diffusion architectures. These findings underscore the critical role of frequency-domain insights for generating robust and reliable motions, paving the way for more intuitive natural language control of robots.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。