音乐大模型的损失值反而因噪音下降,揭示评估新机制
When Noise Lowers The Loss: Rethinking Likelihood-Based Evaluation in Music Large Language Models
- 通过注入不同长度噪音,发现模型对局部干扰更敏感
- 短噪声引发损失骤升,曲线形状比数值更能反映质量
- 适合关注模型评估方法与训练目标优化的研究者
音乐大语言模型的兴起亟需可靠的输出质量评估方法,尤其在区分高质量作品与劣质音乐时。令人困惑的是,标准交叉熵损失——核心训练指标——在遭遇系统性噪声音乐时反而降低,削弱了其作为独立质量指标的有效性。为探究这一矛盾,我们引入噪声注入实验,在音乐上下文中加入可控长度的噪声信号。假设模型对扰动的损失响应,特别是短噪声注入引发的显著上升(“峰值”区域),可作为其辨别音乐完整性的代理指标。在音频波形域的MusicGen模型实验中,结果表明音乐大模型对局部纹理破坏的反应远强于对全局语义破坏的反应。该研究不仅揭示了现有评估的偏差,还提出新原则:损失曲线的形态——而非绝对值——蕴含生成内容质量的关键信息(即模型行为)。我们设想这种基于轮廓的评估可成为无需标签、模型内生的质量评估框架,为更严谨的训练目标和更精准的基准测试开辟道路。
原文摘要 · Abstract (English)
The rise of music large language models (LLMs) demands robust methods of evaluating output quality, especially in distinguishing high-quality compositions from "garbage music". Curiously, we observe that the standard cross-entropy loss -- a core training metric -- often decrease when models encounter systematically corrupted music, undermining its validity as a standalone quality indicator. To investigate this paradox, we introduce noise injection experiment, where controlled noise signal of varying lengths are injected into musical contexts. We hypothesize that a model's loss reacting positively to these perturbations, specifically a sharp increase ("Peak" area) for short injection, can serve as a proxy for its ability to discern musical integrity. Experiments with MusicGen models in the audio waveform domain confirm that Music LLMs respond more strongly to local, texture-level disruptions than to global semantic corruption. Beyond exposing this bias, our results highlight a new principle: the shape of the loss curve -- rather than its absolute value -- encodes critical information about the quality of the generated content (i.e., model behavior). We envision this profile-based evaluation as a label-free, model-intrinsic framework for assessing musical quality -- opening the door to more principled training objectives and sharper benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。