arXiv:2602.18635cs.SDcs.NE2026-02

音乐训练促使神经网络产生音高等价感知,仅听音乐不够。

Musical Training, but not Mere Exposure to Music, Drives the Emergence of Chroma Equivalence in Artificial Neural Networks

  • 用自监督学习听音乐无法生成音高等价,需任务驱动训练。
  • 只有音乐转录任务训练的模型才表现出音高等价特征。
  • 该研究为理解人类音乐感知的形成机制提供新视角。

音高是听觉感知的基本维度,包括音高高度(频率高低感知)和音高等价(八度音程间周期性相似)。现有研究对音高等价是习得还是先天存在存在分歧。本文利用人工神经网络(ANNs)和表示相似性分析,检验了当前听觉模型中音高高度与音高等价的涌现情况。对Wav2Vec 2.0和Data2Vec两个模型进行自监督语音与音乐学习,以及监督式音乐转录任务的微调。结果发现:所有模型均表现出不同程度的音高高度表征,但仅在音乐转录任务训练的模型中观察到音高等价。单纯通过自监督方式接触音乐不足以催生音高等价。这支持音高等价是一种服务于音乐感知的高级认知计算,有别于语音感知。本研究也展示了使用ANNs探索人类感知表征发展条件的有效性。

原文摘要 · Abstract (English)

Pitch is a fundamental aspect of auditory perception. Pitch perception is commonly described across two perceptual dimensions: pitch height is the sense that tones with varying frequencies seem to be higher or lower, and chroma equivalence is the cyclical similarity of notes octaves, corresponding to a doubling of fundamental frequency. Existing research is divided on whether chroma equivalence is a learned percept that varies according to musical experience and culture, or is an innate percept that develops automatically. Building on a recent framework that proposes to use ANNs to ask 'why' questions about the brain, we evaluated recent auditory ANNs using representational similarity analysis to test the emergence of pitch height and chroma equivalence in their learned representations. Additionally, we fine-tuned two models, Wav2Vec 2.0 and Data2Vec, on a self-supervised learning task using speech and music, and a supervised music transcription task. We found that all models exhibited varying degrees of pitch height representation, but that only models trained on the supervised music transcription task exhibited chroma equivalence. Mere exposure to music through self-supervised learning was not sufficient for chroma equivalence to emerge. This supports the view that chroma equivalence is a higher-order cognitive computation that emerges to support the specific task of music perception, distinct from other auditory perception such as speech listening. This work also highlights the usefulness of ANNs for probing the developmental conditions that give rise to perceptual representations in humans.

音高感知神经网络音乐认知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。