arXiv:2601.17844cs.HCcs.AI2026-01

用视觉语言模型分析脑电波图像,实现无需训练的癫痫检测。

RAICL: Retrieval-Augmented In-Context Learning for Vision-Language-Model Based EEG Seizure Detection

  • 将多导脑电信号转为图像,结合领域知识提示,用大模型识别脑波模式。
  • 在少样本条件下动态检索最相关示例,提升模型对非平稳脑电信号的适应性。
  • 不需重新训练,直接使用现成大模型,适合临床快速部署。

脑电图(EEG)解码是医学诊断、康复工程和脑机接口的关键环节。然而,现有解码方法高度依赖特定任务的数据集来训练专用神经网络,数据稀缺限制了通用大型脑解码模型的发展。本文提出一种范式转变:不再基于信号本身解码,而是利用大规模视觉-语言模型(VLMs)分析脑电波形图像。通过将多变量脑电信号转化为堆叠波形图像,并在文本提示中融入神经科学领域知识,我们证明基础VLM可有效区分人脑中的不同活动模式。为应对脑电信号固有的非平稳性,提出检索增强的上下文学习(RAICL)方法,动态选取最具代表性和相关性的少样本示例,以引导VLM的自回归输出。在基于EEG的癫痫检测实验中,采用RAICL的先进VLM表现优于或媲美传统时序方法。结果表明,该方法为生理信号处理开辟了新路径,有效融合视觉、语言与神经活动模态。此外,使用无需再训练或下游架构构建的现成VLM,为临床应用提供了即插即用的解决方案。

原文摘要 · Abstract (English)

Electroencephalogram (EEG) decoding is a critical component of medical diagnostics, rehabilitation engineering, and brain-computer interfaces. However, contemporary decoding methodologies remain heavily dependent on task-specific datasets to train specialized neural network architectures. Consequently, limited data availability impedes the development of generalizable large brain decoding models. In this work, we propose a paradigm shift from conventional signal-based decoding by leveraging large-scale vision-language models (VLMs) to analyze EEG waveform plots. By converting multivariate EEG signals into stacked waveform images and integrating neuroscience domain expertise into textual prompts, we demonstrate that foundational VLMs can effectively differentiate between different patterns in the human brain. To address the inherent non-stationarity of EEG signals, we introduce a Retrieval-Augmented In-Context Learning (RAICL) approach, which dynamically selects the most representative and relevant few-shot examples to condition the autoregressive outputs of the VLM. Experiments on EEG-based seizure detection indicate that state-of-the-art VLMs under RAICL achieved better or comparable performance with traditional time series based approaches. These findings suggest a new direction in physiological signal processing that effectively bridges the modalities of vision, language, and neural activities. Furthermore, the utilization of off-the-shelf VLMs, without the need for retraining or downstream architecture construction, offers a readily deployable solution for clinical applications.

脑电图视觉语言模型少样本学习癫痫检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。