arXiv:2410.22299cs.SDcs.CV2024-10被引 6

根据图片情绪生成匹配音乐,让音乐更懂画面情感。

Emotion-Guided Image to Music Generation

  • 用情绪空间直接控制音乐生成,避免复杂对比学习。
  • 在多个音乐指标上优于现有方法,尤其提升节奏一致性。
  • 适合做图文配乐、社交内容创作的开发者和创作者。

从图像生成音乐可增强幻灯片背景音乐、社交媒体体验及视频制作等应用。本文提出一种情绪引导的图像到音乐生成框架,利用效价-唤醒(Valence-Arousal, VA)情绪空间,使生成音乐与输入图像的情绪基调一致。与以往依赖对比学习保证情绪一致性的方法不同,该方法直接引入VA损失函数,实现更精准的情绪对齐。模型采用CNN-Transformer架构,包含预训练的CNN图像特征提取器和三个Transformer编码器,用于捕捉高阶音乐情绪特征;三个Transformer解码器则对特征进行精炼,生成在音乐性与情绪上一致的MIDI序列。在新构建的图像-MIDI情绪配对数据集上的实验表明,该模型在音符密度(Polyphony Rate)、音高熵(Pitch Entropy)、节奏一致性(Groove Consistency)及损失收敛速度等指标上均表现优异。

原文摘要 · Abstract (English)

Generating music from images can enhance various applications, including background music for photo slideshows, social media experiences, and video creation. This paper presents an emotion-guided image-to-music generation framework that leverages the Valence-Arousal (VA) emotional space to produce music that aligns with the emotional tone of a given image. Unlike previous models that rely on contrastive learning for emotional consistency, the proposed approach directly integrates a VA loss function to enable accurate emotional alignment. The model employs a CNN-Transformer architecture, featuring pre-trained CNN image feature extractors and three Transformer encoders to capture complex, high-level emotional features from MIDI music. Three Transformer decoders refine these features to generate musically and emotionally consistent MIDI sequences. Experimental results on a newly curated emotionally paired image-MIDI dataset demonstrate the proposed model's superior performance across metrics such as Polyphony Rate, Pitch Entropy, Groove Consistency, and loss convergence.

图像音乐生成情绪控制Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。