用深度学习从脑电数据中解码视觉语义,仅需少量样本即可实现高精度预测。
Visual Semantic Decoding of Electrocorticography from Video Stimuli using End-to-End Deep Learning

- 端到端深度学习框架直接从脑电时间序列中预测视频类别,无需人工特征提取。
- 在每类少于50个样本条件下,模型准确率达83.4%,关键依赖高伽马频段(80-150Hz)信号。
- 模型可解释性强,揭示了视觉皮层多个区域对语义解码的重要贡献。
基于脑皮层电图(ECoG)的视觉语义解码能够从复杂且嘈杂的脑活动推断出视觉感知的语义理解。本研究采用端到端深度学习框架,评估从视频刺激中解码视觉类别任务的可行性。使用来自17名药物难治性癫痫患者的数据集,每类视觉类别训练样本不足50个。对比多种深度学习方法、神经网络架构及频率带滤波输入。最优方案结合mixup数据增强、Transformer编码器,使用80-150 Hz高伽马频段、900毫秒后刺激窗口输入。分析显示,早期视觉皮层(V2-V4)、腹侧流视觉皮层、MT+复合区及其邻近区域、以及外侧颞叶对解码性能贡献显著。结果表明,端到端深度学习框架可在无手工特征条件下,实现动态视觉刺激的高效解码,且模型行为在频谱、时间和皮层维度上具有可解释性,与已有神经科学认知高度一致。
原文摘要 · Abstract (English)
ECoG-based visual semantic decoding enables inference of semantic interpretation of visual perception from complex, noisy brain activity. This study examines the feasibility of visual semantic decoding using an end-to-end deep learning framework using electrocorticography (ECoG). Specifically, the decoding task is to predict visual categories from video stimuli using time-series neural inputs. A previously collected ECoG dataset from participants ($n=17$) with drug-resistant epilepsy is used for analysis. With fewer than 50 training samples per visual category, this study evaluates multiple deep learning approaches, artificial neural network architectures, and frequency-band filtered inputs. The best-performing approach is analyzed to shed light on the discriminative information it relies on across spectral, temporal, and cortical dimensions. The selected decoding system uses mixup augmentation, a Transformer-based encoder, and high-gamma (80-150 Hz) inputs with a 900 ms post-stimulus window. Further analysis shows that early visual cortex (V2-V4), ventral stream visual cortex, MT+ complex with neighbouring visual areas, and lateral temporal cortex contributed substantially to decoding performance. This study demonstrates that an end-to-end deep learning framework can yield promising decoding performance from dynamic visual stimuli without handcrafted features, while the model behavior remains interpretable through spectral, temporal, and cortical dimensions, which are broadly consistent with established neuroscience knowledge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。