用Transformer和扩散模型从脑电波重建视觉图像,提升语义理解能力。
Brain-Gen: Towards Interpreting Neural Signals for Stimulus Reconstruction Using Transformers and Latent Diffusion Models
- 基于Transformer提取脑电信号的时空特征,输入扩散模型进行图像生成。
- 在未见类别上零样本泛化能力提升11.8%,潜空间聚类准确率提高6.5%。
- 适合脑机接口、神经机制解析研究者关注,推动可解释性脑信号分析。
神经科学与人工智能的进步使得脑活动解码初见成效。然而,神经表征的可解释性仍受限,主要源于脑电图(EEG)信号固有的高噪声、空间弥散和显著的时间变异性。为揭示思维背后的神经机制,我们提出一种基于Transformer的框架,从EEG记录中提取与视觉刺激相关的时空特征,并将其融入潜空间扩散模型(LDM)的注意力机制,实现从脑活动重建视觉刺激。在公开基准数据集上的定量评估表明,该方法在建模EEG信号语义结构方面表现优异:潜空间聚类准确率最高提升6.5%,零样本泛化能力提升11.8%,同时保持与现有基线相当的Inception Score和Fréchet Inception Distance。本工作标志着迈向可泛化的EEG语义解释的重要一步。
原文摘要 · Abstract (English)
Advances in neuroscience and artificial intelligence have enabled preliminary decoding of brain activity. However, despite the progress, the interpretability of neural representations remains limited. A significant challenge arises from the intrinsic properties of electroencephalography (EEG) signals, including high noise levels, spatial diffusion, and pronounced temporal variability. To interpret the neural mechanism underlying thoughts, we propose a transformers-based framework to extract spatial-temporal representations associated with observed visual stimuli from EEG recordings. These features are subsequently incorporated into the attention mechanisms of Latent Diffusion Models (LDMs) to facilitate the reconstruction of visual stimuli from brain activity. The quantitative evaluations on publicly available benchmark datasets demonstrate that the proposed method excels at modeling the semantic structures from EEG signals; achieving up to 6.5% increase in latent space clustering accuracy and 11.8% increase in zero shot generalization across unseen classes while having comparable Inception Score and Fréchet Inception Distance with existing baselines. Our work marks a significant step towards generalizable semantic interpretation of the EEG signals.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。