arXiv:2510.19332cs.CV2025-10被引 1

用CLIP中间层融合解码脑影像,少参数却更准

BrainMCLIP: Brain Image Decoding with Multi-Layer feature Fusion of CLIP

  • 按人脑视觉层级对齐fMRI信号与CLIP多层特征
  • 高阶语义指标超越现有方法,参数量减少71.7%
  • 无需额外VAE模块,适合追求效率与细节的视觉重建

从fMRI解码图像通常将脑活动映射到CLIP的最终语义层。为捕捉更精细视觉细节,许多方法引入参数密集的VAE流程。但这些方法忽略了CLIP中间层中的丰富对象信息,且违背了大脑功能层次结构。我们提出BrainMCLIP,首次采用参数高效、基于人类视觉系统功能层级的多层特征融合策略,无需额外VAE路径。BrainMCLIP将来自不同功能视觉区域(低/高层)的fMRI信号分别对齐至对应CLIP中间层与最终层,尊重功能层次。我们还引入交叉重构策略和一种新型多粒度损失。结果表明,BrainMCLIP性能极具竞争力,尤其在高阶语义指标上达到或超越现有最先进方法,包括使用VAE流程的方法。关键在于,其参数量相比顶尖VAE基线方法减少71.7%(表 ef{tab:compare_clip_vae}),通过避免VAE路径实现。利用CLIP中间特征,有效捕获常被仅用最终层忽略的视觉细节,在语义准确性和细节保真度间取得良好平衡。

原文摘要 · Abstract (English)

Decoding images from fMRI often involves mapping brain activity to CLIP's final semantic layer. To capture finer visual details, many approaches add a parameter-intensive VAE-based pipeline. However, these approaches overlook rich object information within CLIP's intermediate layers and contradicts the brain's functionally hierarchical. We introduce BrainMCLIP, which pioneers a parameter-efficient, multi-layer fusion approach guided by human visual system's functional hierarchy, eliminating the need for such a separate VAE pathway. BrainMCLIP aligns fMRI signals from functionally distinct visual areas (low-/high-level) to corresponding intermediate and final CLIP layers, respecting functional hierarchy. We further introduce a Cross-Reconstruction strategy and a novel multi-granularity loss. Results show BrainMCLIP achieves highly competitive performance, particularly excelling on high-level semantic metrics where it matches or surpasses SOTA(state-of-the-art) methods, including those using VAE pipelines. Crucially, it achieves this with substantially fewer parameters, demonstrating a reduction of 71.7\%(Table.\ref{tab:compare_clip_vae}) compared to top VAE-based SOTA methods, by avoiding the VAE pathway. By leveraging intermediate CLIP features, it effectively captures visual details often missed by CLIP-only approaches, striking a compelling balance between semantic accuracy and detail fidelity without requiring a separate VAE pipeline.

脑影像解码CLIP融合多层特征轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。