用多尺度视觉编码器解码脑信号,更精准还原视觉信息。
Learning Brain Representation with Hierarchical Visual Embeddings
- 融合多个具不同先验的预训练视觉编码器,捕捉多层次视觉表征。
- 在真实脑电数据上实现高精度图像检索与高质量重建。
- 适合脑机接口、神经科学等领域研究者参考。
从脑信号中解码视觉表征在神经科学与人工智能领域备受关注。然而,脑信号究竟在何种程度上编码了视觉信息仍不明确。现有视觉解码方法虽探索多种脑-图对齐策略,但多侧重高层语义特征,忽视像素级细节,限制了对人类视觉系统的理解。本文提出一种脑-图对齐策略,利用多个具有不同归纳偏置的预训练视觉编码器,捕获分层且多尺度的视觉表征,并采用对比学习目标实现脑信号与视觉嵌入的有效对齐。此外,引入融合先验(Fusion Prior),在大规模视觉数据上学习稳定映射,再将脑特征匹配至该预训练先验,从而增强模态间分布一致性。大量定量与定性实验表明,本方法在图像检索准确率与重建保真度之间取得良好平衡。
原文摘要 · Abstract (English)
Decoding visual representations from brain signals has attracted significant attention in both neuroscience and artificial intelligence. However, the degree to which brain signals truly encode visual information remains unclear. Current visual decoding approaches explore various brain-image alignment strategies, yet most emphasize high-level semantic features while neglecting pixel-level details, thereby limiting our understanding of the human visual system. In this paper, we propose a brain-image alignment strategy that leverages multiple pre-trained visual encoders with distinct inductive biases to capture hierarchical and multi-scale visual representations, while employing a contrastive learning objective to achieve effective alignment between brain signals and visual embeddings. Furthermore, we introduce a Fusion Prior, which learns a stable mapping on large-scale visual data and subsequently matches brain features to this pre-trained prior, thereby enhancing distributional consistency across modalities. Extensive quantitative and qualitative experiments demonstrate that our method achieves a favorable balance between retrieval accuracy and reconstruction fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。