双向控制流让动作生成更自然,支持文本图像多模态风格控制。
MulSMo: Multimodal Stylized Motion Generation by Bidirectional Control Flow
- 构建内容与风格的双向信息流动,缓解风格冲突。
- 在多个数据集上优于现有方法,支持多模态风格输入。
- 适合需要灵活风格控制的动作生成研究者使用。
生成符合目标风格且遵循内容提示的动作序列,需兼顾内容与风格。现有方法通常仅单向传递风格到内容,易导致风格与内容冲突,影响融合效果。本文提出双向控制流机制,使风格能动态适应内容,有效缓解风格-内容冲突,同时更好保留风格动态特性。此外,通过对比学习,将风格控制从单一运动模态扩展至文本、图像等多模态输入,实现更灵活的风格控制。大量实验表明,该方法在多个数据集上显著优于现有方法,并支持多模态信号控制。代码将公开。
原文摘要 · Abstract (English)
Generating motion sequences conforming to a target style while adhering to the given content prompts requires accommodating both the content and style. In existing methods, the information usually only flows from style to content, which may cause conflict between the style and content, harming the integration. Differently, in this work we build a bidirectional control flow between the style and the content, also adjusting the style towards the content, in which case the style-content collision is alleviated and the dynamics of the style is better preserved in the integration. Moreover, we extend the stylized motion generation from one modality, i.e. the style motion, to multiple modalities including texts and images through contrastive learning, leading to flexible style control on the motion generation. Extensive experiments demonstrate that our method significantly outperforms previous methods across different datasets, while also enabling multimodal signals control. The code of our method will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。