arXiv:2605.16923cs.CV2026-05被引 1

模仿人脑视觉处理阶段,分步解码脑电中的视觉信息。

Neuroscience-inspired Staged Representation Learning with Disentangled Coarse- and Fine-Grained Semantics for EEG Visual Decoding

论文配图:Neuroscience-inspired Staged Representation Learning with Disentangled Coarse- and Fine-Grained Semantics for EEG Visual Decoding
图 1 · 摘自论文原文
  • 分三阶段学习低层视觉、高层语义与融合信息,贴近人脑处理机制。
  • 在THINGS-EEG上实现零样本跨被试精准检索,准确率显著提升。
  • 适合脑机接口与神经康复领域研究者,尤其关注跨被试泛化性能。

从脑电信号(EEG)解码视觉信息是脑机接口与医疗康复中的核心挑战。现有方法多聚焦于单一全局嵌入的跨模态对齐,却忽视了人类视觉处理的阶段性与层次性特征。为此,我们提出一种受神经科学启发的分阶段表征学习框架,将EEG视觉解码重构为阶段特异性表征分解问题。该框架包含三个互补阶段:低层视觉表征学习、高层语义表征学习与整合信息融合。为强化语义建模,引入多模态双层级语义学习机制,分离粗粒度标签级语义与细粒度图像级视觉-语义信息。同时,通过从观测到的视觉EEG信号生成语义潜通道,扩展了通道级语义表示空间,实现结构化语义抽象与跨模态对齐。在THINGS-EEG基准上的大量实验表明,所提方法在被试依赖的零样本评估中表现更优,在被试独立的零样本评估中实现更高精确检索。额外分析包括逐层检索、时间累积、多图像扩展检索及消融实验,均支持分阶段分解与结构化语义建模的有效性。结果表明,显式建模感知、语义与整合阶段的表征,是一种有效的神经科学启发型脑电视觉解码框架。

原文摘要 · Abstract (English)

Decoding visual information from electroencephalography (EEG) signals remains a fundamental challenge in brain-computer interfaces and medical rehabilitation. Existing EEG visual decoding methods mainly focus on learning a single global EEG embedding for cross-modal alignment, but they largely overlook the staged and hierarchical characteristics of human visual processing. To address this limitation, we propose a neuroscience-inspired staged representation learning framework that reformulates EEG visual decoding as a stage-specific representation decomposition problem. The proposed framework organizes EEG representation learning into three complementary phases: low-level visual representation learning, high-level semantic representation learning, and integrative information fusion. To strengthen semantic modeling, we further introduce a multimodal dual-level semantic learning mechanism that separates coarse label-level semantics from fine image-level visual-semantic information. In addition, semantic latent channels are introduced as computational representation channels generated from observed visual EEG signals, expanding the channel-level semantic representation space for structured semantic abstraction and cross-modal alignment. Extensive experiments on the THINGS-EEG benchmark demonstrate that the proposed method achieves superior performance under subject-dependent zero-shot evaluation and improved exact retrieval under subject-independent zero-shot evaluation. Additional analyses, including layer-wise retrieval, temporal accumulation, expanded multi-image retrieval, and ablation studies, further support the effectiveness of staged decomposition and structured semantic modeling. These results suggest that explicitly modeling staged perceptual, semantic, and integrative representations provides an effective neuroscience-inspired framework for EEG-based visual decoding.

脑机接口视觉解码分阶段学习神经科学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。