用情绪维度统一图像、音乐和歌词,实现跨模态情感匹配。
MMVA: Multimodal Matching Based on Valence and Arousal across Images, Music, and Musical Captions
- 基于情绪的正负性与强度构建三模态匹配框架
- 在2.5万组图像-音乐对上达到当前最优情感预测效果
- 适用于零样本场景,适合内容生成与推荐系统
我们提出基于情感正负性(valence)与唤醒度(arousal)的多模态匹配框架MMVA,用于捕捉图像、音乐及音乐描述中的情感内容。为支持该框架,我们扩展了Image-Music-Emotion-Matching-Net(IMEMNet)数据集,构建了包含24,756张图像和25,944段音乐片段及其对应音乐描述的IMEMNet-C数据集。采用基于连续情感值的多模态匹配分数,在训练中通过计算不同模态间的正负性与唤醒度相似性,实现随机采样图像-音乐对。所提方法在情感正负性与唤醒度预测任务中达到当前最优性能,并在多种零样本任务中展现有效性,凸显情感预测在下游应用中的潜力。
原文摘要 · Abstract (English)
We introduce Multimodal Matching based on Valence and Arousal (MMVA), a tri-modal encoder framework designed to capture emotional content across images, music, and musical captions. To support this framework, we expand the Image-Music-Emotion-Matching-Net (IMEMNet) dataset, creating IMEMNet-C which includes 24,756 images and 25,944 music clips with corresponding musical captions. We employ multimodal matching scores based on the continuous valence (emotional positivity) and arousal (emotional intensity) values. This continuous matching score allows for random sampling of image-music pairs during training by computing similarity scores from the valence-arousal values across different modalities. Consequently, the proposed approach achieves state-of-the-art performance in valence-arousal prediction tasks. Furthermore, the framework demonstrates its efficacy in various zeroshot tasks, highlighting the potential of valence and arousal predictions in downstream applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。