用多重编码器和弗雷谢音频距离,更客观评估音乐情绪识别与生成。
Addressing Emotion Bias in Music Emotion Recognition and Generation with Frechet Audio Distance
- 结合多种音频编码器与弗雷谢音频距离(FAD),避免单一模型偏差。
- 实验显示真实与合成音乐在情绪表达上存在显著差异,合成音乐情绪较弱。
- 新方法提升生成音乐的情绪多样性和表现力,更接近真实情感。
音乐情绪的复杂性导致在情绪识别(MER)和情绪音乐生成(EMG)中存在固有偏差,尤其当依赖单一音频编码器、情绪分类器或评估指标时。本文对MER和EMG进行研究,采用多种音频编码器并结合无参考评估指标弗雷谢音频距离(FAD)。首先对MER进行基准评估,揭示单一编码器的局限性及不同测量方式间的差异。随后提出基于多编码器FAD的MER性能评估方法,实现更客观的音乐情绪衡量。此外,提出改进的EMG方法,增强生成音乐的情绪多样性与表现力,提升真实性。还比较了真实与合成音乐在情绪表达上的差异,评估所提模型与两个基线模型的表现。实验表明,情绪偏差存在于MER与EMG中,而使用FAD与多编码器可更客观、有效地评估音乐情绪。
原文摘要 · Abstract (English)
The complex nature of musical emotion introduces inherent bias in both recognition and generation, particularly when relying on a single audio encoder, emotion classifier, or evaluation metric. In this work, we conduct a study on Music Emotion Recognition (MER) and Emotional Music Generation (EMG), employing diverse audio encoders alongside Frechet Audio Distance (FAD), a reference-free evaluation metric. Our study begins with a benchmark evaluation of MER, highlighting the limitations of using a single audio encoder and the disparities observed across different measurements. We then propose assessing MER performance using FAD derived from multiple encoders to provide a more objective measure of musical emotion. Furthermore, we introduce an enhanced EMG approach designed to improve both the variability and prominence of generated musical emotion, thereby enhancing its realism. Additionally, we investigate the differences in realism between the emotions conveyed in real and synthetic music, comparing our EMG model against two baseline models. Experimental results underscore the issue of emotion bias in both MER and EMG and demonstrate the potential of using FAD and diverse audio encoders to evaluate musical emotion more objectively and effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。