用脑活动重建图像,兼顾结构与语义,效果更优。
HAVIR: HierArchical Vision to Image Reconstruction using CLIP-Guided Versatile Diffusion
- 双适配器融合拓扑与语义信息,提升重建精度。
- 在复杂场景下仍能准确还原图像结构与含义。
- 适合脑机接口、视觉解码研究者参考。
从脑活动重建视觉信息,连接神经科学与计算机视觉。尽管生成模型已在解码fMRI图像方面取得进展,但高复杂度视觉刺激的精确恢复仍面临挑战,源于其元素密度高、多样性大、空间结构复杂及语义信息丰富。为此,我们提出HAVIR,包含两个适配器:(1) AutoKL适配器将fMRI体素转换为潜在扩散先验,捕捉拓扑结构;(2) CLIP适配器将体素转为CLIP文本与图像嵌入,包含语义信息。二者通过通用扩散模型融合生成最终图像。为提取复杂场景中的关键语义,CLIP适配器使用描述视觉刺激的文本字幕及其合成语义图像进行训练。实验表明,HAVIR在复杂场景下有效恢复视觉刺激的结构特征与语义信息,优于现有模型。
原文摘要 · Abstract (English)
Reconstructing visual information from brain activity bridges the gap between neuroscience and computer vision. Even though progress has been made in decoding images from fMRI using generative models, a challenge remains in accurately recovering highly complex visual stimuli. This difficulty stems from their elemental density and diversity, sophisticated spatial structures, and multifaceted semantic information. To address these challenges, we propose HAVIR that contains two adapters: (1) The AutoKL Adapter transforms fMRI voxels into a latent diffusion prior, capturing topological structures; (2) The CLIP Adapter converts the voxels to CLIP text and image embeddings, containing semantic information. These complementary representations are fused by Versatile Diffusion to generate the final reconstructed image. To extract the most essential semantic information from complex scenarios, the CLIP Adapter is trained with text captions describing the visual stimuli and their corresponding semantic images synthesized from these captions. The experimental results demonstrate that HAVIR effectively reconstructs both structural features and semantic information of visual stimuli even in complex scenarios, outperforming existing models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。