arXiv:2608.29987cs.SDeess.AS2026-08

对比三种生成模型在情绪控制下的表现,发现文本条件生成更稳定,且情绪一致性因类型而异。

How Well Do Generative Music Models Follow Emotion Conditioning?

论文配图:How Well Do Generative Music Models Follow Emotion Conditioning?
图 1 · 摘自论文原文
  • 用文本与音频描述构建情绪提示,统一评估音乐生成的情绪跟随能力
  • 文本生成在情感一致性上优于音频生成,且价态比唤醒度更易保持
  • 不同音乐流派间情绪跟随差异显著,需针对性评估情感可控性

近期生成式音乐模型通过文本和音频条件实现越来越精细的控制,但其对预期情绪线索的忠实程度仍不明确。本文提出统一评估流程,基于GTZAN数据集全部1000首曲目,使用DashengLM提取语义音频描述,并用Music2Emotion模型估计源音频的价态(valence)与唤醒度(arousal)。结合描述与高分情绪标签构建情感感知提示,使用Stable Audio Open、MusicGen和InspireMusic三个系统生成30秒音频,评估文本与音频条件生成效果。通过计算生成音频的价态与唤醒度,与源音频对比绝对误差及价态-唤醒空间欧氏距离衡量情绪跟随性能。结果显示:文本条件生成始终优于音频条件生成,其中MusicGen(文本)与InspireMusic(文本)表现最佳,而音频条件版本稳定性较差;价态比唤醒度更可靠地被保留,且情绪跟随性能在不同音乐流派间差异显著。这些发现强调应直接评估情感可控性,而非仅依赖通用质量或提示相关性指标。

原文摘要 · Abstract (English)

Recent generative music models offer increasingly fine-grained control through text and audio conditioning, yet how faithfully they follow intended emotional cues remains an open question. We address this gap with a unified evaluation pipeline for emotion-following in generated music. Using all 1000 tracks in GTZAN, we extract semantic audio descriptions with DashengLM, an audio captioning model, and estimate source-track valence and arousal with Music2Emotion, a music emotion recognition model. We construct affect-aware text prompts by combining descriptions with top-ranked emotion tags and generate 30-second outputs with three systems, Stable Audio Open, MusicGen, and InspireMusic, evaluating both text- and audio-conditioned generation. To measure emotion-following, we compute valence and arousal on generated audio and compare them with the source tracks using absolute error and Euclidean distance in valence-arousal space. Text-conditioned generation consistently outperforms audio conditioning, with MusicGen (text) and InspireMusic (text) achieving the best performance, while audio-conditioned variants prove less stable. We further find that valence is preserved more reliably than arousal and that emotion-following varies substantially across genres. These findings underscore the importance of evaluating affective controllability directly rather than relying solely on general quality or prompt-relevance metrics.

音乐生成情绪控制生成模型评估基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。