用多模态信息提升脑电视觉重建质量,精度显著提高。
CognitionCapturerPro: Towards High-Fidelity Visual Decoding from EEG/MEG via Multi-modal Information and Asymmetric Alignment
- 融合图像、文本等多模态先验,协同训练增强脑电信号表征
- 在THINGS-EEG数据集上Top-1和Top-5准确率分别提升25.9%和10.6%
- 适合脑机接口、神经解码领域研究者参考
从脑电(EEG)中重构视觉刺激仍面临保真度下降与表征偏移的挑战。我们提出CognitionCapturerPro,一种增强框架,通过协同训练将EEG与多模态先验(图像、文本、深度、边缘)结合。核心贡献包括:基于不确定性的相似性评分机制,用于量化各模态的保真度;以及融合编码器,整合共享表征。采用简化对齐模块与预训练扩散模型,在THINGS-EEG数据集上,相比原版CognitionCapturer,Top-1和Top-5检索准确率分别提升25.9%和10.6%。代码已开源:https://github.com/XiaoZhangYES/CognitionCapturerPro。
原文摘要 · Abstract (English)
Visual stimuli reconstruction from EEG remains challenging due to fidelity loss and representation shift. We propose CognitionCapturerPro, an enhanced framework that integrates EEG with multi-modal priors (images, text, depth, and edges) via collaborative training. Our core contributions include an uncertainty-weighted similarity scoring mechanism to quantify modality-specific fidelity and a fusion encoder for integrating shared representations. By employing a simplified alignment module and a pre-trained diffusion model, our method significantly outperforms the original CognitionCapturer on the THINGS-EEG dataset, improving Top-1 and Top-5 retrieval accuracy by 25.9% and 10.6%, respectively. Code is available at: https://github.com/XiaoZhangYES/CognitionCapturerPro.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。