arXiv:2604.22649cs.NEcs.CV2026-04被引 1

用结构引导扩散模型,从脑电波重建更逼真的视觉图像。

Structure-Guided Diffusion Model for EEG-Based Visual Cognition Reconstruction

论文配图:Structure-Guided Diffusion Model for EEG-Based Visual Cognition Reconstruction
图 1 · 摘自论文原文
  • 通过控制网融合结构信息,指导脑电到图像的生成过程。
  • 在抽象与自然图像上均超越现有方法,细节和语义还原度更高。
  • 适合脑机接口、认知神经科学等领域研究者参考。

目标:从脑电图(EEG)解码视觉信息是神经科学与脑机接口的重要课题。现有方法多局限于自然图像和类别表征,难以捕捉结构特征,也难区分客观感知与主观认知。本文提出结构引导扩散模型(SGDM),引入显式结构信息实现基于EEG的视觉重建。方法:在Kilogram抽象视觉对象数据集和THINGS自然图像数据集上,采用两阶段生成机制。框架结合结构监督变分自编码器与时空脑电信号编码器,通过对比学习对齐视觉嵌入空间;利用ControlNet将结构信息融入扩散模型,引导图像生成。结果:SGDM在两类数据集上均优于现有方法,重建图像在低层视觉特征和语义表征上保真度更高,显示更强解码准确性和跨视觉域泛化能力。脑电信号的时空分析揭示了符合视觉认知神经动态的层次化结构编码模式。意义:验证了SGDM在捕捉显式几何结构与生成高保真认知图像方面的有效性。该框架突破传统低维或类别输出限制,使脑机接口具备更高自由度的意图解码能力,推动脑-机通信向更灵活方向发展。

原文摘要 · Abstract (English)

Objective: Decoding visual information from electroencephalography (EEG) is an important problem in neuroscience and brain-computer interface (BCI) research. Existing methods are largely restricted to natural images and categorical representations, with limited capacity to capture structural features and to differentiate objective perception from subjective cognition. We propose a Structure-Guided Diffusion Model (SGDM) that incorporates explicit structural information for EEG-based visual reconstruction. Approach: SGDM is evaluated on the Kilogram abstract visual object dataset and the THINGS natural image dataset using a two-stage generative mechanism. The framework combines a structurally supervised variational autoencoder with a spatiotemporal EEG encoder aligned to a visual embedding space via contrastive learning. Structural information is integrated into a diffusion model through ControlNet to guide image generation from EEG features. Results: SGDM outperforms existing methods on both abstract and natural image datasets. Reconstructed images achieve higher fidelity in low-level visual features and semantic representations, indicating improved decoding accuracy and strong generalization across diverse visual domains. Spatiotemporal analysis of EEG signals further reveals hierarchical structural encoding patterns, consistent with the neural dynamics of visual cognition. Significance: These findings validate the effectiveness of SGDM in capturing explicit structural geometry and generating images with high fidelity to individual cognitive representations. By enabling decoding of complex visual content from EEG signals, the framework extends neural decoding beyond low-dimensional or categorical outputs. This supports BCIs with increased degrees of freedom for intention decoding and more flexible brain-to-machine communication.

脑机接口图像重建扩散模型结构引导

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。