用大脑对音乐的期待与声音特征区分建模,提升脑电识曲准确率。
Expectation and Acoustic Neural Network Representations Enhance Music Identification from Brain Activity
- 分离声学与预期信息作为神经网络教师信号
- 结合两类表示使识别准确率超越强基线模型
- 无需人工标注即可提取预测性脑响应,适合认知研究
听音乐时,皮层活动同时编码声学信息与预期相关信号。已有研究表明,人工神经网络(ANN)表征与皮层表征相似,可作为脑电(EEG)识别的监督信号。本文证明,将声学与预期相关的ANN表征分别作为教师目标,能显著提升基于EEG的音乐识别性能。预训练模型在预测任一表征时均优于非预训练基线,两者结合带来互补增益,超过通过随机初始化变化形成的强集成模型。结果表明教师表征类型显著影响下游表现,且表征学习可由神经编码引导。本研究为预测性音乐认知和神经解码提供新路径。我们提出的预期表征直接从原始信号计算,无需人工标注,反映超出起始点或音高的预测结构,支持跨多样刺激的多层预测编码研究。其在大规模多样化数据集上的可扩展性,暗示了基于皮层编码原则构建通用型脑电模型的潜力。
原文摘要 · Abstract (English)
During music listening, cortical activity encodes both acoustic and expectation-related information. Prior work has shown that ANN representations resemble cortical representations and can serve as supervisory signals for EEG recognition. Here we show that distinguishing acoustic and expectation-related ANN representations as teacher targets improves EEG-based music identification. Models pretrained to predict either representation outperform non-pretrained baselines, and combining them yields complementary gains that exceed strong seed ensembles formed by varying random initializations. These findings show that teacher representation type shapes downstream performance and that representation learning can be guided by neural encoding. This work points toward advances in predictive music cognition and neural decoding. Our expectation representation, computed directly from raw signals without manual labels, reflects predictive structure beyond onset or pitch, enabling investigation of multilayer predictive encoding across diverse stimuli. Its scalability to large, diverse datasets further suggests potential for developing general-purpose EEG models grounded in cortical encoding principles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。