arXiv:2509.04985cs.SDeess.AS2025-09

提出PAMT模型,让音乐对抗攻击更难被察觉。

Training a Perceptual Model for Evaluating Auditory Similarity in Music Adversarial Attack

  • 在冻结MERT编码器上加轻量投影头,用心理声学条件训练。
  • 与人听觉判断相关性达0.65,对抗攻击下准确率提升9.15%。
  • 适合做音乐识别系统安全评估,尤其对抗攻击研究者。

音乐信息检索(MIR)系统极易受到人类难以察觉的对抗攻击,根源在于模型特征空间与人类听觉感知之间存在错位。现有防御方法和感知度量常无法捕捉这些听觉细微差别,初步听觉测试显示常见度量与人类判断的相关性较低。为此,我们提出感知对齐的MERT Transformer(PAMT),一种学习鲁棒、感知对齐音乐表示的新框架。核心创新在于基于心理声学条件的序列对比变换器,是一个构建在冻结MERT编码器之上的轻量级投影头。PAMT在主观评分上达到0.65的斯皮尔曼相关系数,优于现有感知度量。该方法在覆盖歌曲识别和音乐流派分类等挑战性任务中,平均提升9.15%的鲁棒准确率,且在多种感知对抗攻击下表现优异。本工作首次实现架构内集成的心理声学条件化,生成的表示显著更贴近人类感知,同时增强对音乐对抗攻击的鲁棒性。

原文摘要 · Abstract (English)

Music Information Retrieval (MIR) systems are highly vulnerable to adversarial attacks that are often imperceptible to humans, primarily due to a misalignment between model feature spaces and human auditory perception. Existing defenses and perceptual metrics frequently fail to adequately capture these auditory nuances, a limitation supported by our initial listening tests showing low correlation between common metrics and human judgments. To bridge this gap, we introduce Perceptually-Aligned MERT Transformer (PAMT), a novel framework for learning robust, perceptually-aligned music representations. Our core innovation lies in the psychoacoustically-conditioned sequential contrastive transformer, a lightweight projection head built atop a frozen MERT encoder. PAMT achieves a Spearman correlation coefficient of 0.65 with subjective scores, outperforming existing perceptual metrics. Our approach also achieves an average of 9.15\% improvement in robust accuracy on challenging MIR tasks, including Cover Song Identification and Music Genre Classification, under diverse perceptual adversarial attacks. This work pioneers architecturally-integrated psychoacoustic conditioning, yielding representations significantly more aligned with human perception and robust against music adversarial attacks.

音乐对抗感知对齐心理声学特征表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。