用脑电与注意力图联合生成图像,提升重建质量与视觉一致性
EEG-Driven Image Reconstruction with Saliency-Guided Diffusion Models
- 结合脑电信号与视觉注意力图双重条件控制生成过程
- 在THINGS-EEG数据集上显著提升低/高层特征质量,对齐人眼注意力
- 适合神经解码、医疗诊断与脑机接口等需要高保真图像重建场景
现有脑电驱动图像重建方法常忽略空间注意力机制,限制了重建的清晰度与语义连贯性。为此,我们提出一种双条件框架,将脑电嵌入与空间显著性图联合用于图像生成。方法采用自适应思维映射器(ATM)提取脑电特征,并通过低秩适应(LoRA)微调Stable Diffusion 2.1,实现神经信号与视觉语义的对齐;同时引入ControlNet分支,以显著性图引导生成过程,实现空间控制。在THINGS-EEG数据集上的评估显示,该方法在低层与高层图像特征质量上均优于现有方法,且与人类视觉注意高度一致。结果表明,注意力先验可有效缓解脑电信号模糊性,实现高保真图像重建,适用于医学诊断与神经适应性交互系统,通过高效适配预训练扩散模型推动神经解码进展。
原文摘要 · Abstract (English)
Existing EEG-driven image reconstruction methods often overlook spatial attention mechanisms, limiting fidelity and semantic coherence. To address this, we propose a dual-conditioning framework that combines EEG embeddings with spatial saliency maps to enhance image generation. Our approach leverages the Adaptive Thinking Mapper (ATM) for EEG feature extraction and fine-tunes Stable Diffusion 2.1 via Low-Rank Adaptation (LoRA) to align neural signals with visual semantics, while a ControlNet branch conditions generation on saliency maps for spatial control. Evaluated on THINGS-EEG, our method achieves a significant improvement in the quality of low- and high-level image features over existing approaches. Simultaneously, strongly aligning with human visual attention. The results demonstrate that attentional priors resolve EEG ambiguities, enabling high-fidelity reconstructions with applications in medical diagnostics and neuroadaptive interfaces, advancing neural decoding through efficient adaptation of pre-trained diffusion models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。