用语言引导行为基础模型,实现长序列动作生成。
Plan, Don't Pose: Long Composite Motion Generation with Text-Aligned BFM

- 将语言与预训练行为模型对齐,分步规划动作
- 支持长序列、复合语义描述,生成更稳定可靠
- 适合虚拟角色动画与人机交互场景
文本到动作(T2M)生成在角色动画、虚拟形象和人机交互中有广泛应用。现有方法通常直接从语言生成姿态轨迹或运动标记,要求单一模型同时完成语义理解、长时序结构设计和底层物理实现,导致成本高且在长序列、组合性或语义密集提示下不可靠。我们提出Text2BFM,首个将自然语言与预训练行为基础模型(BFM)对齐的T2M生成框架,不依赖重型端到端运动生成器。Text2BFM在冻结的BFM的潜在策略空间中运行,将其作为可执行的动作先验。通过一个文本对齐的变分行为瓶颈,将BFM策略潜在序列压缩为紧凑的动作表示,兼容语言输入并保留长时序行为结构。生成在该紧凑行为流形上进行,使用轻量级条件生成器,最终解码为驱动预训练冻结BFM的策略潜在变量。通过解耦语义规划与动作执行,Text2BFM实现了高效、鲁棒的T2M生成,在长序列、复合文本描述上表现优异。
原文摘要 · Abstract (English)
Text-to-motion (T2M) generation has broad applications in character animation, virtual avatars, and human-robot interaction. Existing methods typically generate pose trajectories or motion tokens directly from language, forcing a single model to handle semantic interpretation, long-horizon structure, and low-level physical realization. This coupling makes them costly and often unreliable for long, compositional, or semantically dense prompts. We propose Text2BFM, the first framework that aligns natural language with pretrained Behavioral Foundation Models (BFMs) for T2M generation without relying on heavy end-to-end motion generators. Text2BFM operates in the latent policy space of a frozen BFM, using it as an executable motion prior. A text-aligned variational behavioral bottleneck compresses BFM policy-latent sequences into compact motion representations that are compatible with language and preserve long-horizon behavioral structure. Generation is performed in this compact behavioral manifold with a lightweight conditional generator, and the resulting latent encoded behaviors are decoded into policy latents that drive the pretrained frozen BFM. By decoupling semantic planning from motion execution, Text2BFM achieves efficient, robust T2M generation and strong performance on long, compositional textual descriptions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。