arXiv:2412.13237eess.IVcs.CV2024-12被引 1

用非线性网络优化脑影像解码,提升视觉重建的清晰度和语义准确率。

Optimized two-stage AI-based Neural Decoding for Enhanced Visual Stimulus Reconstruction from fMRI Data

  • 设计非线性深度网络替代线性映射,更好捕捉脑活动与视觉表征的复杂关系。
  • 在自然场景数据集上,结构相似性提升约2%,感知相似性提升约4%。
  • 适用于高噪声环境下视觉刺激重建,对结构与语义恢复均有显著改进。

基于AI的神经解码通过生成模型将功能磁共振(fMRI)测量的脑活动映射为潜在层次表示,以重建视觉感知。传统方法使用岭回归线性模型将fMRI转换至潜在空间,再通过预训练变分自编码器(VAE)驱动的潜在扩散模型(LDM)进行解码。由于fMRI数据复杂且含噪,新方法采用两阶段流程:第一阶段生成粗略视觉近似,第二阶段利用带有CLIP嵌入的LDM优化预测。本文提出一种非线性深度网络,用于优化fMRI潜在空间表征,同时调整维度。在自然场景数据集上的实验表明,所提架构相较于基于岭回归的最先进模型,重构图像的结构相似性提升约2%,语义相似性提升约4%。噪声敏感性分析显示,第一阶段对高结构相似性预测至关重要;而高噪声输入虽削弱结构相似性,但对语义影响较小。研究强调了利用非线性关系及两阶段生成式AI在提升噪声fMRI数据下视觉刺激重建保真度中的关键作用。

原文摘要 · Abstract (English)

AI-based neural decoding reconstructs visual perception by leveraging generative models to map brain activity, measured through functional MRI (fMRI), into latent hierarchical representations. Traditionally, ridge linear models transform fMRI into a latent space, which is then decoded using latent diffusion models (LDM) via a pre-trained variational autoencoder (VAE). Due to the complexity and noisiness of fMRI data, newer approaches split the reconstruction into two sequential steps, the first one providing a rough visual approximation, the second on improving the stimulus prediction via LDM endowed by CLIP embeddings. This work proposes a non-linear deep network to improve fMRI latent space representation, optimizing the dimensionality alike. Experiments on the Natural Scenes Dataset showed that the proposed architecture improved the structural similarity of the reconstructed image by about 2\% with respect to the state-of-the-art model, based on ridge linear transform. The reconstructed image's semantics improved by about 4\%, measured by perceptual similarity, with respect to the state-of-the-art. The noise sensitivity analysis of the LDM showed that the role of the first stage was fundamental to predict the stimulus featuring high structural similarity. Conversely, providing a large noise stimulus affected less the semantics of the predicted stimulus, while the structural similarity between the ground truth and predicted stimulus was very poor. The findings underscore the importance of leveraging non-linear relationships between BOLD signal and the latent representation and two-stage generative AI for optimizing the fidelity of reconstructed visual stimuli from noisy fMRI data.

神经解码视觉重建扩散模型fMRI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。