arXiv:2510.03122cs.CVcs.AI2025-10被引 2

用分层结构重建脑活动图像,提升复杂场景还原效果。

HAVIR: HierArchical Vision to Image Reconstruction using CLIP-Guided Versatile Diffusion

论文配图:HAVIR: HierArchical Vision to Image Reconstruction using CLIP-Guided Versatile Diffusion
图 1 · 摘自论文原文
  • 分两阶段提取脑区结构与语义特征,分别生成扩散先验和CLIP嵌入。
  • 在复杂场景下重建质量显著优于现有方法,结构与语义更准确。
  • 适合脑机接口、视觉神经编码研究者参考。

从脑活动重建视觉信息促进了神经科学与计算机视觉的交叉融合。然而,现有方法在恢复高复杂度视觉刺激时仍存在困难,这源于自然场景的特性:低层特征异质性强,高层特征因上下文重叠而语义纠缠。受视觉皮层分层表征理论启发,我们提出HAVIR模型,将视觉皮层分为两个层次区域,分别提取不同特征。结构生成器从空间处理体素中提取结构信息,并转换为潜扩散先验;语义提取器将语义处理体素转换为CLIP嵌入。二者通过通用扩散模型融合,合成最终图像。实验表明,HAVIR在复杂场景下显著提升了重建的结构与语义质量,性能超越现有模型。

原文摘要 · Abstract (English)

The reconstruction of visual information from brain activity fosters interdisciplinary integration between neuroscience and computer vision. However, existing methods still face challenges in accurately recovering highly complex visual stimuli. This difficulty stems from the characteristics of natural scenes: low-level features exhibit heterogeneity, while high-level features show semantic entanglement due to contextual overlaps. Inspired by the hierarchical representation theory of the visual cortex, we propose the HAVIR model, which separates the visual cortex into two hierarchical regions and extracts distinct features from each. Specifically, the Structural Generator extracts structural information from spatial processing voxels and converts it into latent diffusion priors, while the Semantic Extractor converts semantic processing voxels into CLIP embeddings. These components are integrated via the Versatile Diffusion model to synthesize the final image. Experimental results demonstrate that HAVIR enhances both the structural and semantic quality of reconstructions, even in complex scenes, and outperforms existing models.

脑机接口图像重建扩散模型CLIP

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。