arXiv:2510.06988cs.CV2025-10被引 1

仅用文本提示微调扩散模型,无需动捕数据即可生成新动作。

No MoCap Needed: Post-Training Motion Diffusion Models with Reinforcement Learning using Only Textual Prompts

  • 用文本-动作检索网络做奖励信号,强化学习优化生成分布。
  • 跨数据集和留一法实验中,动作质量和多样性显著提升。
  • 无需真实动作数据,适合隐私敏感场景的快速动作迁移。

扩散模型近期在人体动作生成方面取得进展,可从文本提示生成逼真多样的动画。然而,将这些模型适配到未见动作或风格通常需要额外的动作捕捉数据和完整重训练,成本高且难以扩展。本文提出一种基于强化学习的后训练框架,仅使用文本提示对预训练运动扩散模型进行微调,无需任何动作真值数据。该方法利用预训练的文本-动作检索网络作为奖励信号,并通过去噪扩散策略优化(Denoising Diffusion Policy Optimization)优化扩散策略,有效将模型生成分布推向目标领域,而无需成对动作数据。我们在HumanML3D和KIT-ML数据集上评估了跨数据集适应和留一法动作实验,涵盖潜空间与关节空间扩散架构。定量指标与用户研究结果表明,本方法持续提升生成动作的质量与多样性,同时保持原始分布性能。该方法具有灵活性、数据高效性与隐私保护优势,适用于动作迁移任务。

原文摘要 · Abstract (English)

Diffusion models have recently advanced human motion generation, producing realistic and diverse animations from textual prompts. However, adapting these models to unseen actions or styles typically requires additional motion capture data and full retraining, which is costly and difficult to scale. We propose a post-training framework based on Reinforcement Learning that fine-tunes pretrained motion diffusion models using only textual prompts, without requiring any motion ground truth. Our approach employs a pretrained text-motion retrieval network as a reward signal and optimizes the diffusion policy with Denoising Diffusion Policy Optimization, effectively shifting the model's generative distribution toward the target domain without relying on paired motion data. We evaluate our method on cross-dataset adaptation and leave-one-out motion experiments using the HumanML3D and KIT-ML datasets across both latent- and joint-space diffusion architectures. Results from quantitative metrics and user studies show that our approach consistently improves the quality and diversity of generated motions, while preserving performance on the original distribution. Our approach is a flexible, data-efficient, and privacy-preserving solution for motion adaptation.

动作生成扩散模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。