让音乐生成模型实现连续情感控制,突破文本提示的限制。
LARA-Gen: Enabling Continuous Emotion Control for Music Generation Models via Latent Affective Representation Alignment
- 通过隐空间情感对齐,将内部状态与外部理解模型同步。
- 在连续情绪空间中实现精细情感调控,生成质量更高。
- 提供评测基准与预测器,适合音乐创作与情感计算研究者。
近期文本到音乐模型已能根据文本提示生成连贯音乐,但精细情感控制仍存挑战。我们提出 LARA-Gen 框架,通过隐式情感表征对齐(LARA)将内部隐藏状态与外部音乐理解模型对齐,实现有效训练。此外,设计基于连续唤醒-效价空间的情感控制模块,将情感属性从文本内容中解耦,规避了文本提示的瓶颈。同时构建包含精选测试集和稳健情绪预测器的基准,支持对音乐生成中情感可控性的客观评估。大量实验表明,LARA-Gen 实现了连续、细粒度的情感控制,在情感贴合度与音乐质量上显著优于基线模型。生成样本见 https://anonymous2232330.github.io/laragen-web/。
原文摘要 · Abstract (English)
Recent advances in text-to-music models have enabled coherent music generation from text prompts, yet fine-grained emotional control remains unresolved. We introduce LARA-Gen, a framework for continuous emotion control that aligns the internal hidden states with an external music understanding model through Latent Affective Representation Alignment (LARA), enabling effective training. In addition, we design an emotion control module based on a continuous valence-arousal space, disentangling emotional attributes from textual content and bypassing the bottlenecks of text-based prompting. Furthermore, we establish a benchmark with a curated test set and a robust Emotion Predictor, facilitating objective evaluation of emotional controllability in music generation. Extensive experiments demonstrate that LARA-Gen achieves continuous, fine-grained control of emotion and significantly outperforms baselines in both emotion adherence and music quality. Generated samples are available at https://anonymous2232330.github.io/laragen-web/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。