用原型聚类建模同一风格内的多样动作,提升风格化动作生成质量。
ClusterStyle: Modeling Intra-Style Diversity with Prototypical Clustering for Stylized Motion Generation
- 通过原型聚类构建全局与局部风格嵌入空间,捕捉风格内多样性。
- 在多个数据集上优于现有最佳模型,显著提升风格迁移效果。
- 适合研究动作生成、风格控制与个性化运动合成的学者使用。
现有风格化动作生成模型虽能从风格动作中提取特定风格信息并注入内容动作,但难以捕捉同一风格下的多样化运动变化。本文提出基于聚类的框架ClusterStyle,不再对每个风格动作学习无结构嵌入,而是利用一组原型有效建模同一风格类别内多样的动作模式。我们考虑两种风格多样性:同类别风格动作间的全局多样性,以及动作序列时间动态中的局部多样性。这两个组件共同构建全局与局部两个结构化风格嵌入空间,并通过与不可学习的原型锚点对齐进行优化。此外,我们通过风格调制适配器(SMA)增强预训练文本到动作生成模型,以融合风格特征。大量实验表明,本方法在风格化动作生成和动作风格迁移任务中均优于现有最先进模型。
原文摘要 · Abstract (English)
Existing stylized motion generation models have shown their remarkable ability to understand specific style information from the style motion, and insert it into the content motion. However, capturing intra-style diversity, where a single style should correspond to diverse motion variations, remains a significant challenge. In this paper, we propose a clustering-based framework, ClusterStyle, to address this limitation. Instead of learning an unstructured embedding from each style motion, we leverage a set of prototypes to effectively model diverse style patterns across motions belonging to the same style category. We consider two types of style diversity: global-level diversity among style motions of the same category, and local-level diversity within the temporal dynamics of motion sequences. These components jointly shape two structured style embedding spaces, i.e., global and local, optimized via alignment with non-learnable prototype anchors. Furthermore, we augment the pretrained text-to-motion generation model with the Stylistic Modulation Adapter (SMA) to integrate the style features. Extensive experiments demonstrate that our approach outperforms existing state-of-the-art models in stylized motion generation and motion style transfer.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。