用音频压缩模型处理脑电数据,效果超预期。
Adapting Neural Audio Codecs to EEG
- 将音频编码器直接用于脑电信号,仅需预处理适配输入要求。
- 微调后重建精度和泛化能力优于从零训练,保持临床关键信息。
- 多通道扩展模块提升电极间空间依赖建模,适合医疗脑电分析。
脑电(EEG)与音频在采样率、通道结构和尺度上差异显著,但我们发现预训练的神经音频编码器可作为脑电压缩的有效起点,只需对数据进行预处理以满足编码器输入约束。以最先进的音频编码器DAC为基础,我们证明原始脑电可映射到其基于步长的分帧结构,从而直接复用音频预训练的编码器-解码器。无需修改即可实现稳定的脑电重建,且在脑电数据上微调后,性能优于从零训练。通过调整残差码本深度、码本大小和输入采样率,系统性探索了压缩与质量的权衡。为捕捉电极间的空间依赖,提出DAC-MC:一种基于注意力的跨通道聚合与通道特定解码的多通道扩展,同时保留音频预训练初始化。在TUH异常与癫痫数据集上的评估表明,该方法能有效保留临床相关特征,体现在谱图重建误差和下游分类准确率上。
原文摘要 · Abstract (English)
EEG and audio are inherently distinct modalities, differing in sampling rate, channel structure, and scale. Yet, we show that pretrained neural audio codecs can serve as effective starting points for EEG compression, provided that the data are preprocessed to be suitable to the codec's input constraints. Using DAC, a state-of-the-art neural audio codec as our base, we demonstrate that raw EEG can be mapped into the codec's stride-based framing, enabling direct reuse of the audio-pretrained encoder-decoder. Even without modification, this setup yields stable EEG reconstructions, and fine-tuning on EEG data further improves fidelity and generalization compared to training from scratch. We systematically explore compression-quality trade-offs by varying residual codebook depth, codebook (vocabulary) size, and input sampling rate. To capture spatial dependencies across electrodes, we propose DAC-MC, a multi-channel extension with attention-based cross-channel aggregation and channel-specific decoding, while retaining the audio-pretrained initialization. Evaluations on the TUH Abnormal and Epilepsy datasets show that the adapted codecs preserve clinically relevant information, as reflected in spectrogram-based reconstruction loss and downstream classification accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。