用神经放电数据重建高清自然图像,提升视觉解码精度。
SpikeVAEDiff: Neural Spike-based Natural Visual Scene Reconstruction via VD-VAE and Versatile Diffusion
- 分两阶段:先用深度变分自编码器生成低分辨率初图
- 再用扩散模型结合文本/视觉特征优化,实现高保真图像生成
- 特定脑区(如VISI)数据显著提升重建效果,适合脑机接口研究
从神经放电活动中重构自然视觉场景是神经科学与计算机视觉的关键挑战。本文提出SpikeVAEDiff,一种结合超深变分自编码器(VDVAE)与通用扩散模型的两阶段框架,可从神经放电数据生成高分辨率且语义合理的图像。第一阶段,VDVAE将神经放电信号映射到潜在空间,生成低分辨率初步重构;第二阶段,回归模型将放电信号映射至CLIP-Vision与CLIP-Text特征,驱动通用扩散模型通过图像到图像生成完成细节修复。我们在Allen Visual Coding-Neuropixels数据集上评估方法,分析不同脑区表现。结果表明,VISI区域激活最显著,在重构质量中起关键作用。展示成功与失败案例,揭示解码挑战。相比基于fMRI的方法,放电数据具有更高时空分辨率。我们验证了VDVAE有效性,并通过消融实验证明特定脑区数据显著提升性能。
原文摘要 · Abstract (English)
Reconstructing natural visual scenes from neural activity is a key challenge in neuroscience and computer vision. We present SpikeVAEDiff, a novel two-stage framework that combines a Very Deep Variational Autoencoder (VDVAE) and the Versatile Diffusion model to generate high-resolution and semantically meaningful image reconstructions from neural spike data. In the first stage, VDVAE produces low-resolution preliminary reconstructions by mapping neural spike signals to latent representations. In the second stage, regression models map neural spike signals to CLIP-Vision and CLIP-Text features, enabling Versatile Diffusion to refine the images via image-to-image generation. We evaluate our approach on the Allen Visual Coding-Neuropixels dataset and analyze different brain regions. Our results show that the VISI region exhibits the most prominent activation and plays a key role in reconstruction quality. We present both successful and unsuccessful reconstruction examples, reflecting the challenges of decoding neural activity. Compared with fMRI-based approaches, spike data provides superior temporal and spatial resolution. We further validate the effectiveness of the VDVAE model and conduct ablation studies demonstrating that data from specific brain regions significantly enhances reconstruction performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。