arXiv:2508.19527cs.CV2025-08被引 1

用流匹配加速文本驱动动作生成,实现高精度实时动画。

MotionFlux: Efficient Text-Guided Motion Generation through Rectified Flow Matching and Preference Alignment

  • 基于确定性流匹配构建最优传输路径,避免多步采样
  • 新方法在语义一致性和动作质量上超越现有模型
  • 适合需要实时动画的虚拟角色与智能体应用

动作生成对虚拟角色和具身智能体的动画至关重要。尽管近期文本驱动方法取得显著进展,但仍难以精确对齐语言描述与动作语义,且推理效率低下。为此,我们提出TMR++对齐偏好优化(TAPO),通过迭代调整强化语义锚定,使细微动作变化与文本修饰精准对应。为实现实时合成,我们设计MotionFLUX框架,基于确定性修正流匹配构建噪声分布与动作空间间的最优传输路径。相比传统扩散模型需数百次去噪步骤,该方法利用线性化概率路径大幅减少采样步骤,显著提升推理速度。实验表明,TAPO与MotionFLUX协同构成统一系统,在语义一致性、动作质量和生成速度上均优于当前最先进方法。代码与预训练模型将公开。

原文摘要 · Abstract (English)

Motion generation is essential for animating virtual characters and embodied agents. While recent text-driven methods have made significant strides, they often struggle with achieving precise alignment between linguistic descriptions and motion semantics, as well as with the inefficiencies of slow, multi-step inference. To address these issues, we introduce TMR++ Aligned Preference Optimization (TAPO), an innovative framework that aligns subtle motion variations with textual modifiers and incorporates iterative adjustments to reinforce semantic grounding. To further enable real-time synthesis, we propose MotionFLUX, a high-speed generation framework based on deterministic rectified flow matching. Unlike traditional diffusion models, which require hundreds of denoising steps, MotionFLUX constructs optimal transport paths between noise distributions and motion spaces, facilitating real-time synthesis. The linearized probability paths reduce the need for multi-step sampling typical of sequential methods, significantly accelerating inference time without sacrificing motion quality. Experimental results demonstrate that, together, TAPO and MotionFLUX form a unified system that outperforms state-of-the-art approaches in both semantic consistency and motion quality, while also accelerating generation speed. The code and pretrained models will be released.

动作生成流匹配实时合成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。