对比三种生成模型在情绪控制下的表现,发现文本条件生成更稳定,且情绪一致性因类型而异。
How Well Do Generative Music Models Follow Emotion Conditioning?

- 用文本与音频描述构建情绪提示,统一评估音乐生成的情绪跟随能力
- 文本生成在情感一致性上优于音频生成,且价态比唤醒度更易保持
- 不同音乐流派间情绪跟随差异显著,需针对性评估情感可控性
近期生成式音乐模型通过文本和音频条件实现越来越精细的控制,但其对预期情绪线索的忠实程度仍不明确。本文提出统一评估流程,基于GTZAN数据集全部1000首曲目,使用DashengLM提取语义音频描述,并用Music2Emotion模型估计源音频的价态(valence)与唤醒度(arousal)。结合描述与高分情绪标签构建情感感知提示,使用Stable Audio Open、MusicGen和InspireMusic三个系统生成30秒音频,评估文本与音频条件生成效果。通过计算生成音频的价态与唤醒度,与源音频对比绝对误差及价态-唤醒空间欧氏距离衡量情绪跟随性能。结果显示:文本条件生成始终优于音频条件生成,其中MusicGen(文本)与InspireMusic(文本)表现最佳,而音频条件版本稳定性较差;价态比唤醒度更可靠地被保留,且情绪跟随性能在不同音乐流派间差异显著。这些发现强调应直接评估情感可控性,而非仅依赖通用质量或提示相关性指标。
原文摘要 · Abstract (English)
Recent generative music models offer increasingly fine-grained control through text and audio conditioning, yet how faithfully they follow intended emotional cues remains an open question. We address this gap with a unified evaluation pipeline for emotion-following in generated music. Using all 1000 tracks in GTZAN, we extract semantic audio descriptions with DashengLM, an audio captioning model, and estimate source-track valence and arousal with Music2Emotion, a music emotion recognition model. We construct affect-aware text prompts by combining descriptions with top-ranked emotion tags and generate 30-second outputs with three systems, Stable Audio Open, MusicGen, and InspireMusic, evaluating both text- and audio-conditioned generation. To measure emotion-following, we compute valence and arousal on generated audio and compare them with the source tracks using absolute error and Euclidean distance in valence-arousal space. Text-conditioned generation consistently outperforms audio conditioning, with MusicGen (text) and InspireMusic (text) achieving the best performance, while audio-conditioned variants prove less stable. We further find that valence is preserved more reliably than arousal and that emotion-following varies substantially across genres. These findings underscore the importance of evaluating affective controllability directly rather than relying solely on general quality or prompt-relevance metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。