用波形模板字典实现低数据量下可解释的脑电分析
Interpretable EEG biomarkers with bag-of-waves: Spatial and temporal waveform dictionaries for low-data regimes

- 无监督学习波形原子,通过平移不变k均值构建字典
- 在仅16只小鼠数据上实现基因型分类,性能媲美深度模型
- 原子可直接可视化,适合临床神经生理学家验证
脑电图(EEG)广泛用于神经系统疾病诊断,但其分析通常依赖预设频谱特征或深度神经网络。预设特征带有强先验偏差,而深度模型难以解释且需大量数据和算力。本文提出bag-of-waves框架,通过无标签的平移不变k均值学习少量重复出现的脑电波形模板(称为原子),将连续脑电信号转化为原子标记序列,再用于下游分类或聚类。方法扩展为引入原子间转移关系(n-gram)以捕捉时序结构,并从单通道原子推广至区域及跨通道空间原子。在三个互补数据集上验证:仅16只小鼠的单通道基因型聚类(低数据与时序场景)、静息态痴呆分类(空间场景)、TUEV基准任务(六分类临床事件,高数据对比)。结果表明,bag-of-waves在所有任务中表现媲美前沿深度与基础模型,参数量仅为后者的几分之一,且具备完全可解释性——每个原子对应可观测波形,能显式恢复已知临床形态,供神经生理学家直接验证。其核心优势在于低数据场景下的有效性,传统重型模型在此不适用。
原文摘要 · Abstract (English)
Electroencephalography (EEG) is widely used to diagnose neurological conditions, but its analysis usually relies on either predefined spectral features or deep neural networks. Predefined features carry a strong bias, since they fix in advance what counts as informative, while deep neural networks and foundation models are hard to interpret and need large amounts of data and compute. We present bag-of-waves, an interpretable framework that learns a small dictionary of recurring EEG waveform templates, called atoms, using shift-invariant k-means without labels. The continuous EEG is then turned into a sequence of atom tokens, whose counts feed a simple downstream classifier or clustering step. We extend this representation in two ways: we add atom-to-atom transitions, which we call n- grams, to capture temporal structure, and we move from single-channel atoms to regional and cross-channel spatial atoms for the multichannel case. We test the method on three complementary datasets, each probing a different aspect: single-channel mouse genotype clustering with only sixteen animals (the low-data and temporal case), resting-state dementia classification (the spatial case), and the TUEV benchmark, a six-way classification of clinical EEG events (a high-data comparison against strong deep and foundation baselines). Across all three datasets, bag-of-waves achieves performance competitive with state-of-the-art deep and foundation models. Yet, it operates with a fraction of the parameter count and provides full interpretability: because every atom corresponds to an inspectable waveform, the method explicitly recovers known clinical morphologies that a neurophysiologist can directly validate. Its main advantage is that it works in the low-data regime where heavier models are a poor fit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。