arXiv:2605.30925cs.CVcs.GR2026-05International Conf…

让文字生成动作更完整,同时处理多个动作描述。

MultiAct: Text-to-Motion Generation from Composite Text via Tailored Attention Guidance

论文配图:MultiAct: Text-to-Motion Generation from Composite Text via Tailored Attention Guidance
图 1 · 摘自论文原文
  • 通过动态增强注意力,让被忽略的动作成分得到更多关注。
  • 在复合指令下,语义覆盖度提升,动作仍保持真实自然。
  • 无需重训练,适配现有动作生成模型,适合动画与交互应用。

文本到动作生成近年来发展迅速,为动画和人机交互提供了富有表现力的接口。然而,当前模型在处理同时包含多个动作的复合描述时仍显脆弱:往往只聚焦于主导动作,忽略其他成分,导致运动结果不完整或模糊。本文提出 MultiAct,一种无需配对数据、仅在推理阶段运行的框架,可直接作用于预训练的动作生成器,无需重新训练或修改架构。该方法通过自适应放大与被忽视提示成分相关的交叉注意力分数,缓解语义坍缩问题。我们发现有效的调制依赖于提示特定的选择(如目标词元和层),因此引入轻量级辅助决策机制,自动选择最优的注意力增强方式。大量定量与定性评估表明,MultiAct 在复合提示上持续优于现有基线,在提升语义覆盖度的同时保持动作的真实性。项目页面:https://natsala13.github.io/multiact.github.io。

原文摘要 · Abstract (English)

Text-to-motion generation has progressed rapidly in recent years, offering an expressive interface for animation and human-computer interaction. However, current models remain brittle when handling prompts that describe multiple actions occurring at the same time. Rather than realizing all components of a composite description, models frequently prioritize a single dominant action and neglect the rest, leading to incomplete or ambiguous motion. We present MultiAct, an unpaired, inference-time framework for compositional text-to-motion synthesis that operates directly on pretrained motion generators without retraining or architectural modification. Our method counteracts semantic collapse by adaptively amplifying cross-attention scores associated with underrepresented prompt components. We note that effective modulation depends on prompt-specific choices, such as which tokens and layers to target, and introduce a lightweight auxiliary decision scheme that determines the most effective attention-strengthening parametrization. Extensive quantitative and qualitative evaluations demonstrate that MultiAct consistently outperforms existing baselines on composite prompts, achieving improved semantic coverage while preserving motion realism. Project page: https://natsala13.github.io/multiact.github.io.

动作生成文本到动作注意力机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。