用表情和动作控制音乐生成,让旋律更贴合画面。
EXPOTION: Facial Expression and Motion Control for Multimodal Music Generation
- 用面部表情和上半身动作作为音乐生成的控制信号。
- 在7小时同步数据集上实现音画精准对齐,优于现有模型。
- 适合做交互式音乐创作或影视配乐的开发者使用。
我们提出Expotion(面部表情与动作控制的多模态音乐生成模型),通过人类面部表情、上半身动作及文本提示生成富有表现力且时间准确的音乐。在预训练文本到音乐生成模型上采用参数高效微调(PEFT),仅用小规模数据即可实现对多模态控制的精细适配。为确保视频与音乐的精确同步,引入时间平滑策略以对齐多种模态。实验表明,结合视觉特征与文本描述后,生成音乐在音乐性、创造力、节拍-速度一致性、与视频的时间对齐度以及文本遵循度等方面均显著提升,超越所提基线与现有最先进的视频到音乐生成模型。此外,我们构建了一个包含7小时同步视频记录的新数据集,涵盖与音乐对应的丰富面部及上半身动作,为多模态与交互式音乐生成研究提供了重要资源。
原文摘要 · Abstract (English)
We propose Expotion (Facial Expression and Motion Control for Multimodal Music Generation), a generative model leveraging multimodal visual controls - specifically, human facial expressions and upper-body motion - as well as text prompts to produce expressive and temporally accurate music. We adopt parameter-efficient fine-tuning (PEFT) on the pretrained text-to-music generation model, enabling fine-grained adaptation to the multimodal controls using a small dataset. To ensure precise synchronization between video and music, we introduce a temporal smoothing strategy to align multiple modalities. Experiments demonstrate that integrating visual features alongside textual descriptions enhances the overall quality of generated music in terms of musicality, creativity, beat-tempo consistency, temporal alignment with the video, and text adherence, surpassing both proposed baselines and existing state-of-the-art video-to-music generation models. Additionally, we introduce a novel dataset consisting of 7 hours of synchronized video recordings capturing expressive facial and upper-body gestures aligned with corresponding music, providing significant potential for future research in multimodal and interactive music generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。