arXiv:2508.03780cs.SDcs.AI2025-08

可解释模型在音乐情绪识别中更抗干扰,效果接近对抗训练但更快。

Are Inherently Interpretable Models More Robust? A Study In Music Emotion Recognition

  • 用可解释模型聚焦有意义特征,提升对无关扰动的鲁棒性。
  • 可解释模型在对抗样本下表现优于黑箱模型,媲美对抗训练模型。
  • 无需复杂训练,计算成本更低,适合追求效率的场景。

深度学习模型应能对感知相似的新样本产生相似输出,这种能力称为鲁棒性。然而,现有模型易受微小(对抗性)扰动影响,导致输出剧烈变化并暴露对虚假相关性的依赖。本文研究了内在可解释的深度模型是否比黑箱模型更鲁棒。通过对比可解释模型、黑箱模型和对抗训练模型在音乐情绪识别任务中的表现,发现可解释模型在面对对抗样本时表现更稳定,其鲁棒性可达到对抗训练模型水平,且计算开销更低。

原文摘要 · Abstract (English)

One of the desired key properties of deep learning models is the ability to generalise to unseen samples. When provided with new samples that are (perceptually) similar to one or more training samples, deep learning models are expected to produce correspondingly similar outputs. Models that succeed in predicting similar outputs for similar inputs are often called robust. Deep learning models, on the other hand, have been shown to be highly vulnerable to minor (adversarial) perturbations of the input, which manage to drastically change a model's output and simultaneously expose its reliance on spurious correlations. In this work, we investigate whether inherently interpretable deep models, i.e., deep models that were designed to focus more on meaningful and interpretable features, are more robust to irrelevant perturbations in the data, compared to their black-box counterparts. We test our hypothesis by comparing the robustness of an interpretable and a black-box music emotion recognition (MER) model when challenged with adversarial examples. Furthermore, we include an adversarially trained model, which is optimised to be more robust, in the comparison. Our results indicate that inherently more interpretable models can indeed be more robust than their black-box counterparts, and achieve similar levels of robustness as adversarially trained models, at lower computational cost.

可解释性鲁棒性音乐情绪识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。