用轻量级方法实现文本驱动动作生成的风格化,无需为每种风格重训练模型。
Stylized Text-to-Motion Generation via Hypernetwork-Driven Low-Rank Adaptation

- 通过超网络生成低秩参数,动态调节预训练扩散模型的风格
- 在HumanML3D和100STYLE数据集上达到当前最佳风格化效果
- 支持未见过的风格泛化,适合需要灵活风格控制的研究者
文本驱动的动作扩散模型能生成逼真的动作,但仅靠文本难以表达动作的细微风格特征。现有方法要么需为每种风格微调模型,要么依赖重型ControlNet结构,限制了效率与对未知风格的泛化能力。本文提出一种轻量级风格条件框架,通过超网络生成的LoRA参数动态调节预训练扩散模型。将风格参考动作编码为全局风格嵌入,由超网络映射为每个去噪步骤的低秩更新。通过监督对比损失构建风格潜在空间,有效捕捉多样风格属性,提升对未知风格的泛化能力,并支持无需预定义类别优化引导。在HumanML3D和100STYLE数据集上的实验表明,该方法达到当前最优风格化性能,同时显著提升对未见风格的生成效果。
原文摘要 · Abstract (English)
Text-driven motion diffusion models are capable of generating realistic human motions, but text alone often struggles to express fine-level nuances of motion, commonly referred to as style. Recent approaches have tackled this challenge by attaching a style injection mechanism to a pretrained text-driven diffusion model. Existing stylization methods, however, either require style-specific fine-tuning of existing models or rely on heavy ControlNet-based architectures, limiting efficiency and generalization to unseen styles. We propose a lightweight style conditioning framework that dynamically modulates a pretrained diffusion model through hypernetwork-generated LoRA parameters. A style reference motion is encoded into a global style embedding, which is mapped by a hypernetwork to low-rank updates applied at each denoising step of the diffusion model. By structuring the style latent space with a supervised contrastive loss, our framework reliably captures diverse stylistic attributes, improves generalization to unseen styles, and supports optimization-based guidance without requiring predefined style categories. Experiments on the HumanML3D and 100STYLE datasets show state-of-the-art stylization results, while achieving improved stylization for unseen styles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。