用指令微调和反馈对齐让音乐大模型学会预测情绪高低
Aligning MusicLLM with Emotion using Instruction Tuning and Feedback-Driven Alignment

- 通过任务感知指令微调让模型学习情绪回归
- 反馈驱动对齐使唤醒度与效价预测准确率显著提升
- 在保持问答能力的同时提升情绪预测性能
本文研究音乐大语言模型(MusicLLMs)在情绪回归任务上的对齐可能性。尽管MusicLLMs在音乐信息检索任务中表现优异,但其对唤醒度和效价的预测能力仍受限,因情绪回归未作为明确训练目标。为检验MusicLLMs能否实现情绪对齐,我们分别采用指令微调与反馈驱动对齐两种策略进行训练。实验表明,任务感知的指令微调使MusicLLMs具备一定情绪水平预测能力,但准确率有限;而结合可验证数值奖励的反馈驱动对齐,显著提升了唤醒度与效价的预测表现。此外,该方法在提升情绪回归性能的同时,保持了MusicQA任务的能力。
原文摘要 · Abstract (English)
This paper investigates whether music large language models (MusicLLMs) can be aligned for emotion regression. While MusicLLMs have shown strong performance in music information retrieval tasks, their ability to predict arousal and valence scores remains limited, since emotion regression has not been an explicit training objective. To examine whether MusicLLMs can be aligned with emotion, we train MusicLLMs on emotion regression and compare two strategies: instruction tuning and feedback-driven alignment. Our experiments show that task-aware instruction tuning enables MusicLLMs to predict emotion levels to some extent, although the accuracy remains limited. Applying feedback-driven alignment with a verifiable numerical reward substantially improves performance on both arousal and valence over instruction tuning alone. We further show that our approach improves emotion regression performance while maintaining MusicQA capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。