用生物启发的自监督方法,让脑电图更准地还原图像。
NeuroBridge: Bio-Inspired Self-Supervised EEG-to-Image Decoding via Cognitive Priors and Bidirectional Semantic Alignment

- 引入认知先验和双向语义对齐,增强脑电与图像的跨模态匹配。
- 在200类零样本检索任务中,准确率提升至63.2%(顶1)和89.9%(顶5)。
- 适用于脑机接口与认知研究,尤其适合缺乏标注数据的场景。
视觉神经解码旨在从脑活动模式中重建或推断感知的视觉刺激,为理解人类认知提供关键洞见,并推动脑机接口与人工智能的变革性应用。然而,现有方法受限于高质量刺激-脑反应配对数据稀缺,以及神经表征与视觉内容之间的语义错位。受生物系统感知变异性和协同适应策略启发,我们提出一种新型自监督架构NeuroBridge,融合认知先验增强(CPA)与共享语义投影器(SSP),以实现有效跨模态对齐。具体而言,CPA通过施加非对称、模态特定的变换模拟感知变异,提升语义多样性;与以往方法不同,SSP采用协同适应策略,实现双模态特征在共享语义空间中的双向对齐,促进有效跨模态学习。NeuroBridge在同被试与跨被试设置下均超越现有最先进方法。在同被试场景中,顶1准确率提升12.3%,顶5准确率提升10.2%,在200类零样本检索任务中分别达到63.2%与89.9%。大量实验验证了该框架在神经视觉解码中的有效性、鲁棒性与可扩展性。
原文摘要 · Abstract (English)
Visual neural decoding seeks to reconstruct or infer perceived visual stimuli from brain activity patterns, providing critical insights into human cognition and enabling transformative applications in brain-computer interfaces and artificial intelligence. Current approaches, however, remain constrained by the scarcity of high-quality stimulus-brain response pairs and the inherent semantic mismatch between neural representations and visual content. Inspired by perceptual variability and co-adaptive strategy of the biological systems, we propose a novel self-supervised architecture, named NeuroBridge, which integrates Cognitive Prior Augmentation (CPA) with Shared Semantic Projector (SSP) to promote effective cross-modality alignment. Specifically, CPA simulates perceptual variability by applying asymmetric, modality-specific transformations to both EEG signals and images, enhancing semantic diversity. Unlike previous approaches, SSP establishes a bidirectional alignment process through a co-adaptive strategy, which mutually aligns features from two modalities into a shared semantic space for effective cross-modal learning. NeuroBridge surpasses previous state-of-the-art methods under both intra-subject and inter-subject settings. In the intra-subject scenario, it achieves the improvements of 12.3% in top-1 accuracy and 10.2% in top-5 accuracy, reaching 63.2% and 89.9% respectively on a 200-way zero-shot retrieval task. Extensive experiments demonstrate the effectiveness, robustness, and scalability of the proposed framework for neural visual decoding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。