arXiv:2502.02063cs.CVcs.AI2025-02被引 3

让文字更精准控制动作生成,提升真实感和可控性

CASIM: Composite Aware Semantic Injection for Text to Motion Generation

  • 用动态对应关系替代固定长度文本嵌入,捕捉动作复合特性
  • 在HumanML3D和KIT数据集上显著提升动作质量与对齐度
  • 适配自回归与扩散模型,对新文本有更强泛化能力

生成建模与分词技术的进步推动了文本到动作生成的显著进展,提升了生成动作的质量与真实性。然而,如何有效利用文本信息进行条件化动作生成仍是挑战。现有方法主要依赖固定长度文本嵌入(如CLIP)进行全局语义注入,难以捕捉人体动作的复合特性,导致动作质量与可控性不足。为此,我们提出复合感知语义注入机制(CASIM),包括复合感知语义编码器与文本-动作对齐器,学习文本与动作标记之间的动态对应关系。CASIM具有模型与表示无关性,可无缝集成于自回归与扩散基方法。在HumanML3D和KIT基准上的实验表明,CASIM在多种前沿方法中持续提升动作质量、文本-动作对齐度与检索得分。定性分析进一步显示,相比固定长度语义注入,本方法能实现更精确的动作控制,并对未见文本输入具备更强泛化能力。

原文摘要 · Abstract (English)

Recent advances in generative modeling and tokenization have driven significant progress in text-to-motion generation, leading to enhanced quality and realism in generated motions. However, effectively leveraging textual information for conditional motion generation remains an open challenge. We observe that current approaches, primarily relying on fixed-length text embeddings (e.g., CLIP) for global semantic injection, struggle to capture the composite nature of human motion, resulting in suboptimal motion quality and controllability. To address this limitation, we propose the Composite Aware Semantic Injection Mechanism (CASIM), comprising a composite-aware semantic encoder and a text-motion aligner that learns the dynamic correspondence between text and motion tokens. Notably, CASIM is model and representation-agnostic, readily integrating with both autoregressive and diffusion-based methods. Experiments on HumanML3D and KIT benchmarks demonstrate that CASIM consistently improves motion quality, text-motion alignment, and retrieval scores across state-of-the-art methods. Qualitative analyses further highlight the superiority of our composite-aware approach over fixed-length semantic injection, enabling precise motion control from text prompts and stronger generalization to unseen text inputs.

动作生成文本控制语义对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。