arXiv:2503.21775cs.CVcs.AI2025-03ICCV被引 7

用多模态输入生成风格化动作,支持文本、图像、音频等风格迁移。

StyleMotif: Multi-Modal Motion Stylization using Style-Content Cross Fusion

  • 通过跨模态融合机制,联合建模动作内容与风格特征。
  • 在多个数据集上优于现有方法,生成动作更逼真且风格一致。
  • 适合需要多模态风格控制的动作生成任务,如动画与虚拟人。

我们提出StyleMotif,一种新型的风格化运动潜空间扩散模型,可基于多模态输入(包括运动序列、文本、图像、视频和音频)生成同时包含内容与风格的动作。不同于仅关注动作多样性或单向风格迁移的方法,StyleMotif 能在广泛的内容基础上无缝融合多模态风格线索。为此,我们设计了风格-内容交叉融合机制,并将风格编码器与预训练的多模态模型对齐,确保生成动作准确捕捉参考风格的同时保持真实感。大量实验表明,该框架在风格化动作生成上超越现有方法,并展现出多模态动作风格化的涌现能力,实现更细腻的动作合成。代码与预训练模型将在论文录用后公开。项目页面:https://stylemotif.github.io

原文摘要 · Abstract (English)

We present StyleMotif, a novel Stylized Motion Latent Diffusion model, generating motion conditioned on both content and style from multiple modalities. Unlike existing approaches that either focus on generating diverse motion content or transferring style from sequences, StyleMotif seamlessly synthesizes motion across a wide range of content while incorporating stylistic cues from multi-modal inputs, including motion, text, image, video, and audio. To achieve this, we introduce a style-content cross fusion mechanism and align a style encoder with a pre-trained multi-modal model, ensuring that the generated motion accurately captures the reference style while preserving realism. Extensive experiments demonstrate that our framework surpasses existing methods in stylized motion generation and exhibits emergent capabilities for multi-modal motion stylization, enabling more nuanced motion synthesis. Source code and pre-trained models will be released upon acceptance. Project Page: https://stylemotif.github.io

动作生成多模态风格迁移扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。