arXiv:2510.02266cs.CVcs.HC2025-10

用轻量框架实现跨人脑高精度视觉重建,仅需1小时训练

NeuroSwift: A Lightweight Cross-Subject Framework for fMRI Visual Reconstruction of Complex Scenes

  • 融合扩散模型与CLIP,分层重建低级特征与语义信息
  • 仅微调17%参数,跨被试重建效果达最新水平
  • 单个4090显卡一小时完成训练,适合资源受限场景

通过计算机视觉技术从脑活动重建视觉信息,有助于直观理解视觉神经机制。尽管生成模型在解码fMRI数据方面取得进展,但实现复杂视觉输入的跨被试准确重建仍具挑战性,且计算成本高。这主要源于个体间神经表征差异以及大脑对复杂视觉输入核心语义特征的抽象编码。为此,我们提出NeuroSwift,通过扩散模型整合互补适配器:AutoKL用于低级特征,CLIP用于语义。其CLIP适配器在Stable Diffusion生成图像与COCO描述配对数据上训练,模拟高级视觉皮层编码。为实现跨被试泛化,先在一个被试上预训练,再仅微调17%参数(全连接层)即可用于新被试,其余组件冻结。该方法在轻量级GPU(三块RTX 4090)上每被试仅需一小时训练,性能超越现有方法。

原文摘要 · Abstract (English)

Reconstructing visual information from brain activity via computer vision technology provides an intuitive understanding of visual neural mechanisms. Despite progress in decoding fMRI data with generative models, achieving accurate cross-subject reconstruction of visual stimuli remains challenging and computationally demanding. This difficulty arises from inter-subject variability in neural representations and the brain's abstract encoding of core semantic features in complex visual inputs. To address these challenges, we propose NeuroSwift, which integrates complementary adapters via diffusion: AutoKL for low-level features and CLIP for semantics. NeuroSwift's CLIP Adapter is trained on Stable Diffusion generated images paired with COCO captions to emulate higher visual cortex encoding. For cross-subject generalization, we pretrain on one subject and then fine-tune only 17 percent of parameters (fully connected layers) for new subjects, while freezing other components. This enables state-of-the-art performance with only one hour of training per subject on lightweight GPUs (three RTX 4090), and it outperforms existing methods.

脑机接口视觉重建轻量模型跨被试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。