用文字和动作参考生成更灵活的人体动画,风格更细腻。
Flexible Motion Generation from Language and Style References

- 用文本+动作片段联合控制生成,风格更精准。
- 支持长序列多风格变化,比以往方法更灵活。
- 无需标注风格标签,能泛化到未见过的组合。
我们提出 FlexMoGen,一种基于自然语言描述和动作风格参考的灵活人体运动生成框架。文本提示能定义语义内容,但难以捕捉时间节奏、肢体动作细节和表现力动态等细微风格特征;而风格参考片段可直接传递这些细节,使模型在保持高层意图的同时复现目标风格。给定文本提示和风格参考片段,FlexMoGen 生成高质量运动,兼具内容准确性和风格一致性。与依赖离散风格标签、无法支持长时序或多风格生成的旧方法不同,FlexMoGen 采用无监督风格编码器,支持长时间、时变、多风格合成。其在统一架构中联合预训练风格编码器与文本到运动的潜在扩散模型,通过轻量级适配模块调节风格。结合高效的相对位置编码,模型在风格化与非风格化数据集上训练,具备对未见文本-风格组合的强大泛化能力。实验表明,FlexMoGen 在内容保真度与风格还原之间取得最佳平衡。
原文摘要 · Abstract (English)
We introduce FlexMoGen, a novel framework for flexible human motion synthesis conditioned on both natural language descriptions and motion style references. Text prompts are effective at defining semantic content, but they are often limited in capturing fine-grained style details such as timing, limb articulation, and expressive dynamics. A style example clip supplements the text by conveying these nuanced motion characteristics directly, enabling the model to preserve high-level intent while reproducing the desired stylistic traits. Given a text prompt and a style example clip, FlexMoGen generates high-quality motions that preserve semantic content while faithfully reflecting the target style, offering users greater control over the animation generation process. Unlike prior methods that rely on discrete style labels and do not generalize to long or multi-style generation, FlexMoGen learns a variational style encoder without style supervision and supports long, time-varying, multi-style synthesis. Our framework jointly pre-trains the style encoder and a text-to-motion latent diffusion model within a unified architecture, modulating motion style through a lightweight adaptation module. It integrates an efficient relative positional encoding scheme and is trained on both stylized and non-stylized datasets, enabling strong generalization to unseen text-style combinations. Experiments show that FlexMoGen achieves the best balance between content fidelity and style reflection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。