首个评估文本生成音乐情绪传达能力的基准,揭示模型情绪偏差。
AImoclips: A Benchmark for Evaluating Emotion Conveyance in Text-to-Music Generation
- 构建涵盖12种情绪的基准AImoclips,测试1000+音乐片段
- 商业模型偏于愉悦,开源模型偏于平淡,高唤醒情绪更准确
- 发现所有模型都倾向情绪中性,适合情感可控研究者参考
文本到音乐(TTM)生成技术虽能实现可控且富有表现力的音乐创作,但其情感保真度仍远未被充分探索,相比人类偏好或文本对齐而言。本文提出AImoclips,一个评估TTM系统向听者传达预期情绪能力的基准,覆盖开源与商业模型。我们选取了跨越效价-唤醒四象限的12种情绪意图,使用六种先进TTM系统生成超过1,000段音乐片段。共111名参与者在9点李克特量表上对每段音乐的感知效价与唤醒度进行评分。结果表明:商业系统生成的音乐普遍比预期更愉悦,而开源系统则相反;所有模型在高唤醒条件下情绪传达更准确;此外,所有系统均表现出对情绪中性的偏差,凸显情感可控性的关键局限。该基准为理解模型特定的情绪渲染特性提供了宝贵洞见,支持未来情感对齐的TTM系统研发。
原文摘要 · Abstract (English)
Recent advances in text-to-music (TTM) generation have enabled controllable and expressive music creation using natural language prompts. However, the emotional fidelity of TTM systems remains largely underexplored compared to human preference or text alignment. In this study, we introduce AImoclips, a benchmark for evaluating how well TTM systems convey intended emotions to human listeners, covering both open-source and commercial models. We selected 12 emotion intents spanning four quadrants of the valence-arousal space, and used six state-of-the-art TTM systems to generate over 1,000 music clips. A total of 111 participants rated the perceived valence and arousal of each clip on a 9-point Likert scale. Our results show that commercial systems tend to produce music perceived as more pleasant than intended, while open-source systems tend to perform the opposite. Emotions are more accurately conveyed under high-arousal conditions across all models. Additionally, all systems exhibit a bias toward emotional neutrality, highlighting a key limitation in affective controllability. This benchmark offers valuable insights into model-specific emotion rendering characteristics and supports future development of emotionally aligned TTM systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。