arXiv:2511.05350cs.SDcs.AI2025-11中稿 · EUSIPCO 2026

用噪声自编码器让音乐表征更贴近人耳感知层次。

Perceptually Aligning Representations of Music via Noise-Augmented Autoencoders

  • 在编码噪声后重建,引入感知损失函数。
  • 粗粒度表征能更好捕捉听觉显著信息。
  • 提升音高意外感与脑电响应预测效果。

我们提出,将自编码器训练为从编码的噪声版本中重构输入,并结合感知驱动的损失函数,可生成符合感知层级结构的表征。通过实验表明,采用此方法训练音频自编码器后,感知显著信息被更高效地保留在较粗糙的表示层级中,相较传统训练方式更为优越。进一步验证显示,这种感知层级结构有助于改进潜在扩散模型在音乐中音高意外感估计和听众脑电反应预测任务中的表现,优于已有方法。预训练权重已发布于 github.com/CPJKU/pa-audioic。

原文摘要 · Abstract (English)

We argue that training autoencoders to reconstruct inputs from noised versions of their encodings, when combined with perceptually motivated losses, yields encodings that are structured according to a perceptual hierarchy. We demonstrate the emergence of this hierarchy by showing that, after training an audio autoencoder in this manner, perceptually salient information is captured in coarser representation structures than with conventional training. Furthermore, we show that such perceptual hierarchies improve latent diffusion decoding in the context of estimating pitch surprisal in music and predicting EEG-brain responses to music listening. In both cases, our results surpass those of previous methods. Pretrained weights are available on github.com/CPJKU/pa-audioic.

音乐表征自编码器感知建模扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。