arXiv:2511.17547eess.SPcs.AI2025-11被引 1

用轻量适配器+冻结编码器,让脑电图生成高清图像更准确。

SYNAPSE: Synergizing an Adapter and Finetuning for High-Fidelity EEG Synthesis from a CLIP-Aligned Encoder

  • 先用CLIP对齐的自编码器学语义化脑电信号表示
  • 在CVPR40数据集上达到顶尖图像质量与重建效率
  • 适合脑机接口、神经科学领域研究者参考

基于扩散模型的生成技术已实现多模态高质量图像合成。将其扩展至脑信号可深化对人类感知与心理表征的理解。然而,脑电图(EEG)因高噪声、低空间分辨率及强烈个体差异,难以直接应用。现有方法如DreamDiffusion、BrainVis和GWIT依赖复杂的对齐或分类流程,将EEG特征适配到预训练的Stable Diffusion模型,导致参数量大且可解释性差。本文提出SYNAPSE,一种两阶段框架:第一阶段,通过结合信号重建与跨模态对齐目标,使CLIP对齐的EEG自编码器学习语义结构化的潜在表示;第二阶段,冻结预训练编码器,仅引入轻量级适配模块接入Stable Diffusion,实现高效条件生成且仅需极少可训练参数。该方法在CVPR40数据集上达成语义连贯的潜在空间与最先进的感知保真度,优于以往EEG-to-image模型,在重建效率与图像质量上均表现更优。定量与定性分析表明,SYNAPSE在跨被试间具有强泛化能力,即使类别层面一致性下降,仍能保留视觉语义。结果表明,还原大脑所感知的内容而非所分类的信息,是实现真实脑电图像生成的关键。

原文摘要 · Abstract (English)

Recent progress in diffusion-based generative models has enabled high-quality image synthesis conditioned on diverse modalities. Extending such models to brain signals could deepen our understanding of human perception and mental representations. However,electroencephalography (EEG) presents major challenges for image generation due to high noise, low spatial resolution, and strong inter-subject variability. Existing approaches,such as DreamDiffusion, BrainVis, and GWIT, primarily adapt EEG features to pre-trained Stable Diffusion models using complex alignment or classification pipelines, often resulting in large parameter counts and limited interpretability. We introduce SYNAPSE, a two-stage framework that bridges EEG signal representation learning and high-fidelity image synthesis. In Stage1, a CLIP-aligned EEG autoencoder learns a semantically structured latent representation by combining signal reconstruction and cross-modal alignment objectives. In Stage2, the pretrained encoder is frozen and integrated with a lightweight adaptation of Stable Diffusion, enabling efficient conditioning on EEG features with minimal trainable parameters. Our method achieves a semantically coherent latent space and state-of-the-art perceptual fidelity on the CVPR40 dataset, outperforming prior EEG-to-image models in both reconstruction efficiency and image quality. Quantitative and qualitative analyses demonstrate that SYNAPSE generalizes effectively across subjects, preserving visual semantics even when class-level agreement is reduced. These results suggest that reconstructing what the brain perceives, rather than what it classifies, is key to faithful EEG-based image generation.

脑电生成扩散模型轻量适配

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。