用脑启发的提示调优,让脑电图更准匹配图像。
NeuroCLIP: Brain-Inspired Prompt Tuning for EEG-to-Image Multimodal Contrastive Learning
- 设计双流视觉嵌入,动态生成图像级提示,精准调控视觉表征。
- 引入全局视觉提示词,在零样本图像检索中达63.2%准确率,提升12.3%。
- 融合神经科学原理优化对比损失,适合跨被试通用的脑机接口研究。
近期脑启发的人工智能致力于通过多模态模型(如CLIP)对齐神经信号与视觉语义。然而,现有方法常将CLIP视为静态特征提取器,忽视其对神经表征的适应性及脑电信号与图像间的生理-符号鸿沟。为此,我们提出NeuroCLIP,一种面向脑电图到图像对比学习的提示调优框架。该方法包含三项核心创新:(1) 设计双流视觉嵌入管道,结合动态滤波与标记级融合,生成实例级自适应提示,引导基于图像内容调整块嵌入标记,实现神经约束下的细粒度视觉表征调制;(2) 首次在脑电图-图像对齐中引入视觉提示标记,作为全局模态级提示,与实例级调节协同工作。这些提示标记插入Transformer架构,促进神经感知适配与全局参数优化;(3) 受人类视觉编码神经科学原理启发,提出改进的对比损失函数,更好建模脑电信号中的语义模糊性与跨模态噪声。在THINGS-EEG2数据集上,NeuroCLIP实现零样本图像检索63.2%的Top-1准确率,较此前最优方法提升+12.3%,并在跨被试条件下表现更强泛化能力(+4.6% Top-1),凸显生理感知提示调优在弥合脑信号与视觉语义方面的潜力。
原文摘要 · Abstract (English)
Recent advances in brain-inspired artificial intelligence have sought to align neural signals with visual semantics using multimodal models such as CLIP. However, existing methods often treat CLIP as a static feature extractor, overlooking its adaptability to neural representations and the inherent physiological-symbolic gap in EEG-image alignment. To address these challenges, we present NeuroCLIP, a prompt tuning framework tailored for EEG-to-image contrastive learning. Our approach introduces three core innovations: (1) We design a dual-stream visual embedding pipeline that combines dynamic filtering and token-level fusion to generate instance-level adaptive prompts, which guide the adjustment of patch embedding tokens based on image content, thereby enabling fine-grained modulation of visual representations under neural constraints; (2) We are the first to introduce visual prompt tokens into EEG-image alignment, acting as global, modality-level prompts that work in conjunction with instance-level adjustments. These visual prompt tokens are inserted into the Transformer architecture to facilitate neural-aware adaptation and parameter optimization at a global level; (3) Inspired by neuroscientific principles of human visual encoding, we propose a refined contrastive loss that better model the semantic ambiguity and cross-modal noise present in EEG signals. On the THINGS-EEG2 dataset, NeuroCLIP achieves a Top-1 accuracy of 63.2% in zero-shot image retrieval, surpassing the previous best method by +12.3%, and demonstrates strong generalization under inter-subject conditions (+4.6% Top-1), highlighting the potential of physiology-aware prompt tuning for bridging brain signals and visual semantics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。