arXiv:2502.16060cs.LGcs.AI2025-02中稿 · ICLR被引 14

用时频模式学习实现单通道脑电的高效分词,提升模型性能与泛化能力

Tokenizing Single-Channel EEG with Time-Frequency Motif Learning

  • 设计双路径时频掩码架构,从单通道脑电信号中学习时频模式词汇
  • 在4个数据集上最高提升11%的科恩卡帕系数,耳部脑电睡眠分期任务提升14%
  • 无需特定设备布局,适配多种基础模型,尤其适合便携式脑电应用

基础模型正在重塑脑电分析,但单通道脑电分词仍是难题。本文提出TFM-Tokenizer,一种从单通道脑电信号中学习时频模式词汇并编码为离散标记的新框架。采用时频掩码的双路径结构捕捉鲁棒的模式表示,且模型无关,可兼容轻量Transformer及现有基础模型。实验显示三大优势:准确率方面,在四个多样化脑电基准测试中,无论单数据集或跨数据集预训练均持续提升,较强基线最高提升11%的科恩卡帕系数;泛化性方面,作为即插即用组件,显著提升BIOT和LaBraM等基础模型表现;可扩展性方面,不依赖10-20标准电极系统,具备设备无关潜力。在耳部脑电睡眠分期任务中,因信号格式、电极配置、设备与任务均不同于预训练数据,仍比基线高出14%。全面的标记分析揭示其具有强类别区分性、频率感知性和结构一致性,提升了表征质量与可解释性。代码已开源:https://github.com/Jathurshan0330/TFM-Tokenizer。

原文摘要 · Abstract (English)

Foundation models are reshaping EEG analysis, yet an important problem of EEG tokenization remains a challenge. This paper presents TFM-Tokenizer, a novel tokenization framework that learns a vocabulary of time-frequency motifs from single-channel EEG signals and encodes them into discrete tokens. We propose a dual-path architecture with time-frequency masking to capture robust motif representations, and it is model-agnostic, supporting both lightweight transformers and existing foundation models for downstream tasks. Our study demonstrates three key benefits: Accuracy: Experiments on four diverse EEG benchmarks demonstrate consistent performance gains across both single- and multi-dataset pretraining settings, achieving up to $11\%$ improvement in Cohen's Kappa over strong baselines. Generalization: Moreover, as a plug-and-play component, it consistently boosts the performance of diverse foundation models, including BIOT and LaBraM. Scalability: By operating at the single-channel level rather than relying on the strict 10-20 EEG system, our method has the potential to be device-agnostic. Experiments on ear-EEG sleep staging, which differs from the pretraining data in signal format, channel configuration, recording device, and task, show that our tokenizer outperforms baselines by $14\%$. A comprehensive token analysis reveals strong class-discriminative, frequency-aware, and consistent structure, enabling improved representation quality and interpretability. Code is available at https://github.com/Jathurshan0330/TFM-Tokenizer.

脑电分析时频模式分词方法单通道EEG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。