用小模型实现跨人情绪解码,精度超大模型。
Dive into Waves: Morlet Spectral Transformer for Cross-Subject Emotion Decoding from EEG

- 用莫莱特小波分频处理脑电信号,匹配大脑节律结构。
- 无需预训练,在多个数据集上超越大模型性能。
- 可解释性强,适合医疗与脑机接口场景使用。
我们研究从脑电图(EEG)中进行跨被试情绪识别,这是脑机接口中一个实际重要但极具挑战的问题。与具有明显波形特征的任务不同,情绪相关脑电信号主要编码在频谱功率中,且信号微弱、噪声大、个体间差异显著。现有方法要么依赖大规模预训练脑电基础模型,需海量数据却仍难应对跨人变异;要么采用频域编码器,虽能反映频谱结构,但存在表征不匹配、漂移主导的分词和缺乏频带特定空间建模等问题。本文提出莫莱特频谱变压器(Morlet Spectral Transformer, MST),基于三个核心组件并融合时空变压器主干网络:首先,莫莱特小波分词提供匹配脑电多尺度节律的时间-频率表示,并将经典微分熵特征扩展为适配变压器的形式;其次,长时基线去除作为简单的时间归一化,消除个体特异性漂移及邻近窗口冗余;第三,频带特定空间投影为每个频带学习独立通道混合器,捕捉可解释的频带特异性模式并减少跨通道混叠。实验表明,即使无预训练,MST在所有SEED系列数据集上均持续优于大型预训练脑电模型和基于频域的方法。结果表明,精心设计的表征可提供一种准确、低成本且可解释的替代方案,无需大规模预训练。
原文摘要 · Abstract (English)
We study cross-subject emotion recognition from EEG, a practically important yet challenging problem in brain-computer interfaces. Unlike tasks with clear waveform signatures, emotion-related EEG signals are primarily encoded in spectral power and are weak, noisy, and highly variable across subjects. Existing approaches rely either on large pretrained EEG foundation models, which require massive data yet still struggle with cross-subject variability, or frequency-domain encoders, which better reflect spectral structure but suffer from mismatched representations, drift-dominated tokenization, and lack of band-specific spatial modeling. In this article, we propose the Morlet Spectral Transformer (MST), built around three key components and integrated with a spatiotemporal Transformer backbone. First, Morlet wavelet tokenization provides a time-frequency representation that matches the multi-scale structure of brain rhythms, and extends classical differential entropy features to a form suitable for Transformers. Second, long-context baseline removal acts as a simple temporal normalization that removes subject-specific drift and redundancy across nearby windows. Third, frequency-specific spatial projection learns a separate channel mixer for each frequency band, capturing interpretable band-specific patterns and reducing cross-channel mixing. We show that, even without pretraining, MST consistently outperforms both large pretrained EEG foundation models and frequency-based methods across all SEED-family datasets. These results suggest that careful representation design can yield an accurate, cost-effective, and interpretable alternative to large-scale pretraining.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。