arXiv:2608.21445cs.CV2026-08

用视觉语言对齐提升脑电图癫痫检测泛化能力

ViTexSZ: Heterogeneous Vision-Text Knowledge Distillation for EEG Seizure Detection

论文配图:ViTexSZ: Heterogeneous Vision-Text Knowledge Distillation for EEG Seizure Detection
图 1 · 摘自论文原文
  • 将脑电图转为图像,通过查询机制对齐不同通道特征
  • 融合临床提示与脑电特征,实现跨模态知识迁移
  • 轻量学生模型在4个数据集上准确率最高,提升达12.9%

从脑电图(EEG)中自动检测癫痫发作对于持续神经监测至关重要,尤其适用于仅表现出细微电生理变化的亚临床癫痫发作。现有时间序列方法通常针对固定电极布局设计,限制了其在通道布局不规则的异构脑电记录中的适用性。尽管视觉与语言建模提供了有前景的替代方案,但将异构脑电表示与临床语义对齐仍具挑战。本文提出ViTexSZ,一种用于脑电图癫痫检测的异构视觉-文本知识蒸馏框架。ViTexSZ将脑电记录转换为结构化波形图像,并引入基于查询的多通道对齐模块,将源相关视觉特征映射至统一标记空间。异构教师模型通过多模态大语言模型将对齐后的脑电表示与临床提示融合,将高层临床语义与发作相关证据关联。视觉-文本知识蒸馏随后将教师表示传递给轻量级学生模型以完成检测。在四个脑电图癫痫数据集上的实验表明,ViTexSZ在亚临床及一般癫痫检测场景下均具有优异泛化能力,在所有数据集上取得最高准确率,相对次优基线提升高达12.9%,验证了其有效性。

原文摘要 · Abstract (English)

Automated seizure detection from electroencephalography (EEG) is essential for continuous neurological monitoring, particularly for subclinical epileptic seizures that may exhibit only subtle electrographic changes. Existing time-series methods are often designed for fixed EEG channel configurations, thereby limiting their applicability to heterogeneous EEG recordings with irregular channel layouts. Although visual and language modeling offer promising alternatives, aligning heterogeneous EEG representations with clinical semantics remains challenging. We introduce ViTexSZ, a heterogeneous Vision-Text knowledge distillation framework for EEG seizure detection. ViTexSZ converts EEG recordings into structured waveform images and introduces a query-based multi-channel alignment module that maps source-dependent visual features into a unified token space. A heterogeneous teacher further integrates the aligned EEG representations with clinical prompts through a multimodal large language model, associating high-level clinical semantics with seizure-related evidence. Vision-text knowledge distillation then transfers the teacher representations to a lightweight student during detection. Experiments on four EEG seizure datasets demonstrate the generalizability of ViTexSZ across both subclinical and general seizure detection scenarios, achieving the highest accuracy on all datasets and relative improvements of up to 12.9% over the second-best baselines, showing its effectiveness.

脑电图分析知识蒸馏多模态学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。