arXiv:2512.00194cs.CVcs.LG2025-12

用视觉语言模型自动识别脑电图伪迹,像专家一样看图说话。

AutocleanEEG ICVision: Automated ICA Artifact Classification Using Vision-Language AI

  • 用大模型直接分析脑电图的多种可视化图表,模仿专家判断
  • 在3168个成分上达成0.677的专家共识一致率,优于传统方法
  • 输出带解释的分类结果,适合需要可解释性的脑电研究者

我们提出EEG Autoclean Vision Language AI(ICVision),首个通过视觉-语言AI实现专家级脑电图独立成分分析(ICA)成分分类的系统。不同于依赖手工特征的传统分类器(如ICLabel),ICVision直接解析ICA仪表盘中的拓扑图、时间序列、功率谱和事件相关电位图,利用多模态大语言模型(GPT-4 Vision)进行视觉理解与自然语言推理。该系统将每个成分归类为六种典型类别之一(脑源、眼动、心跳、肌电、通道噪声、其他噪声),并返回置信度评分与人类可读的解释。在124个脑电数据集共3,168个成分上评估,其与专家共识的一致性达到k = 0.677,超过MNE ICLabel;在模糊案例中仍能保留临床相关脑信号。超过97%的输出被专家评为可理解且可操作。作为开源EEG Autoclean平台的核心模块,ICVision标志着科学人工智能范式转变:模型不仅分类,还能‘看见’、‘思考’并‘沟通’。它开启了全球可扩展、可解释、可复现的脑电分析新路径,推动具备专家级视觉决策能力的AI代理在脑科学乃至更广泛领域的诞生。

原文摘要 · Abstract (English)

We introduce EEG Autoclean Vision Language AI (ICVision) a first-of-its-kind system that emulates expert-level EEG ICA component classification through AI-agent vision and natural language reasoning. Unlike conventional classifiers such as ICLabel, which rely on handcrafted features, ICVision directly interprets ICA dashboard visualizations topography, time series, power spectra, and ERP plots, using a multimodal large language model (GPT-4 Vision). This allows the AI to see and explain EEG components the way trained neurologists do, making it the first scientific implementation of AI-agent visual cognition in neurophysiology. ICVision classifies each component into one of six canonical categories (brain, eye, heart, muscle, channel noise, and other noise), returning both a confidence score and a human-like explanation. Evaluated on 3,168 ICA components from 124 EEG datasets, ICVision achieved k = 0.677 agreement with expert consensus, surpassing MNE ICLabel, while also preserving clinically relevant brain signals in ambiguous cases. Over 97% of its outputs were rated as interpretable and actionable by expert reviewers. As a core module of the open-source EEG Autoclean platform, ICVision signals a paradigm shift in scientific AI, where models do not just classify, but see, reason, and communicate. It opens the door to globally scalable, explainable, and reproducible EEG workflows, marking the emergence of AI agents capable of expert-level visual decision-making in brain science and beyond.

脑电图AI代理视觉语言模型可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。