arXiv:2604.02973cs.CV2026-04

让文字描述与动作更对齐,生成更自然的人体运动。

Exploring Motion-Language Alignment for Text-driven Motion Generation

  • 用全局运动先验+局部文本条件,提升动作与文字匹配度。
  • 发现注意力过度集中在句首,导致关键信息被忽略。
  • 提出新度量和调控策略,适合做动作生成的开发者参考。

文本驱动的人体动作生成旨在根据文字描述合成真实动作序列。尽管近期取得进展,准确对齐动作动态与文本语义仍是核心挑战。本文从动作-语言对齐视角重新审视该问题,提出MLA-Gen框架,融合全局运动先验与细粒度局部条件,使模型既能捕捉常见动作模式,又能建立文本与动作的精细对应关系。此外,我们发现人类动作生成中存在此前未被关注的注意力集中现象:注意力过度聚焦于句子起始标记,削弱了重要语义线索的利用,导致语义定位能力下降。为此,我们引入SinkRatio指标衡量注意力集中程度,并设计对齐感知的掩码与控制策略,在生成过程中调节注意力分布。大量实验表明,所提方法在动作质量与动作-语言对齐性上均优于强基线。代码将在录用后开源。

原文摘要 · Abstract (English)

Text-driven human motion generation aims to synthesize realistic motion sequences that follow textual descriptions. Despite recent advances, accurately aligning motion dynamics with textual semantics remains a fundamental challenge. In this paper, we revisit text-to-motion generation from the perspective of motion-language alignment and propose MLA-Gen, a framework that integrates global motion priors with fine-grained local conditioning. This design enables the model to capture common motion patterns, while establishing detailed alignment between texts and motions. Furthermore, we identify a previously overlooked attention sink phenomenon in human motion generation, where attention disproportionately concentrates on the start text token, limiting the utilization of informative textual cues and leading to degraded semantic grounding. To analyze this issue, we introduce SinkRatio, a metric for measuring attention concentration, and develop alignment-aware masking and control strategies to regulate attention during generation. Extensive experiments demonstrate that our approach consistently improves both motion quality and motion-language alignment over strong baselines. Code will be released upon acceptance.

动作生成文本对齐注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。