用语言模型生成可调节情绪的音乐疗法,帮助患者缓解焦虑与抑郁。
Language Models for Music Medicine Generation
- 基于MusicGen模型,通过低秩微调生成符合情绪过渡规律的30秒音乐片段。
- 在MTG-Jamendo数据集上按情绪标签生成音乐,经情感识别模型验证有效性。
- 合成15分钟音乐疗程,适合作为专业治疗间的临时情绪调节工具。
近年来,音乐疗法被证实对情绪健康具有多重益处。维持良好情绪状态对帕金森病患者或压力、焦虑患者等治疗过程尤为有效。本文提出对MusicGen这一音乐生成Transformer模型进行微调,生成短时音乐片段,辅助患者从负面情绪过渡到目标情绪状态。基于MTG-Jamendo数据集,采用低秩分解微调方法,结合情绪标签生成符合情感维度(效价-唤醒度环形模型)的30秒音乐片段。通过音乐情绪识别模型评估生成音乐的情感一致性。将多个片段拼接后形成一段15分钟的“音乐医学”播放序列,模拟一次音乐治疗过程。本方法是首个利用语言模型生成音乐医学内容的工作。最终输出旨在作为专业音乐治疗师之间短暂干预的补充手段。
原文摘要 · Abstract (English)
Music therapy has been shown in recent years to provide multiple health benefits related to emotional wellness. In turn, maintaining a healthy emotional state has proven to be effective for patients undergoing treatment, such as Parkinson's patients or patients suffering from stress and anxiety. We propose fine-tuning MusicGen, a music-generating transformer model, to create short musical clips that assist patients in transitioning from negative to desired emotional states. Using low-rank decomposition fine-tuning on the MTG-Jamendo Dataset with emotion tags, we generate 30-second clips that adhere to the iso principle, guiding patients through intermediate states in the valence-arousal circumplex. The generated music is evaluated using a music emotion recognition model to ensure alignment with intended emotions. By concatenating these clips, we produce a 15-minute "music medicine" resembling a music therapy session. Our approach is the first model to leverage Language Models to generate music medicine. Ultimately, the output is intended to be used as a temporary relief between music therapy sessions with a board-certified therapist.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。