arXiv:2603.27040cs.CV2026-03

无需预设人数,一次生成多角色动作,效率更高更准确。

Unified Number-Free Text-to-Motion Generation Via Flow Matching

  • 用统一潜空间融合不同数据集,支持跨场景动作生成。
  • 分阶段生成:先产前导动作,再逐轮调整反应,减少误差累积。
  • 适合需要灵活多人动作生成的场景,如游戏、影视制作。

生成模型在固定人数动作合成上表现优异,但在可变人数场景下泛化能力差。现有方法依赖自回归模型递归生成,存在效率低和误差累积问题。本文提出统一运动流(UMF),包含金字塔运动流(P-Flow)与半噪声运动流(S-Flow)。UMF将无定数动作生成分解为单次前导动作生成与多次反应生成阶段。通过统一潜空间弥合异构运动数据分布差距,实现有效联合训练。P-Flow在分层分辨率上基于不同噪声水平生成,降低计算开销;S-Flow学习联合概率路径,自适应完成反应转换与上下文重构,缓解误差积累。大量实验与用户研究验证了UMF作为通用多角色动作生成模型的有效性。

原文摘要 · Abstract (English)

Generative models excel at motion synthesis for a fixed number of agents but struggle to generalize with variable agents. Based on limited, domain-specific data, existing methods employ autoregressive models to generate motion recursively, which suffer from inefficiency and error accumulation. We propose Unified Motion Flow (UMF), which consists of Pyramid Motion Flow (P-Flow) and Semi-Noise Motion Flow (S-Flow). UMF decomposes the number-free motion generation into a single-pass motion prior generation stage and multi-pass reaction generation stages. Specifically, UMF utilizes a unified latent space to bridge the distribution gap between heterogeneous motion datasets, enabling effective unified training. For motion prior generation, P-Flow operates on hierarchical resolutions conditioned on different noise levels, thereby mitigating computational overheads. For reaction generation, S-Flow learns a joint probabilistic path that adaptively performs reaction transformation and context reconstruction, alleviating error accumulation. Extensive results and user studies demonstrate UMF' s effectiveness as a generalist model for multi-person motion generation from text. Project page: https://githubhgh.github.io/umf/.

动作生成文本生成扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。