让脑电模型适应不同电极布局,提升通用性
CAMEL-CLIP: Channel-aware Multimodal Electroencephalography-text Alignment for Generalizable Brain Foundation Models
- 用通道语义信息做位置编码,识别不同电极
- 动态投影各通道,不压缩特征,适配任意数量电极
- 双层对比学习,兼顾单通道和整体信号特征
脑电图(EEG)基础模型在学习通用表征方面展现出潜力,但仍对通道异质性敏感,例如电极数量或排列的变化。我们提出通道感知的多模态脑电-文本对齐对比语言图像预训练模型(CAMEL-CLIP),一种对异构通道配置具有鲁棒性的对比式脑电-文本多模态基础模型,适用于多种下游任务。CAMEL-CLIP引入三个关键组件:(1) 基于通道属性的位置编码,通过语义信息识别通道;(2) 动态通道投影,独立投影每个通道并生成可变长度嵌入,无需特征压缩;(3) 双层对比学习,联合进行通道级与样本级对比学习,以捕捉通道特异性与全局信号特征。实验结果表明,CAMEL-CLIP 在线性探测下达到当前最佳性能,并优于依赖全微调的现有基础模型。
原文摘要 · Abstract (English)
Electroencephalography (EEG) foundation models have shown promise for learning generalizable representations, yet they remain sensitive to channel heterogeneity, such as changes in channel composition or ordering. We propose channel-aware multimodal EEG-text alignment contrastive language-image pretraining (CAMEL-CLIP), a contrastive EEG-text multimodal foundation model designed to be robust to heterogeneous channel configurations and widely applicable to diverse downstream tasks. CAMEL-CLIP introduces three key components: (1) channel attribute-based positional encoding, which identifies channels through semantic information; (2) dynamic channel projection, which generates variable-length embeddings by independently projecting each channel without feature compression; and (3) dual-level contrastive learning, which jointly performs channel-level and sample-level contrastive learning to capture both channel-specific and global signal characteristics. Experimental results demonstrate that CAMEL-CLIP achieves state-of-the-art performance under linear-probing and outperforms existing foundation models that rely on full-finetuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。