arXiv:2604.26218cs.CV2026-04

用视觉图像生成脑电/磁信号,提升脑机接口还原精度。

ViBE: Visual-to-M/EEG Brain Encoding via Spatio-Temporal VAE and Distribution-Aligned Projection

  • 设计时空卷积VAE重建脑电信号动态特征
  • 通过对比学习映射图像特征到脑信号空间
  • 结合误差与分布对齐,实现跨模态精准匹配

脑编码模型不仅有助于理解视觉刺激如何转化为神经反应,也是实现视觉假体以恢复严重视力障碍患者视觉的关键步骤。该过程包含两个核心环节:精确重建神经反应,以及在视觉刺激与神经反应之间建立跨模态对齐。为此,我们提出ViBE,一种从视觉刺激生成脑电(EEG)和脑磁(MEG)信号的新框架。首先,设计时空卷积变分自编码器(TSC-VAE),捕捉M/EEG信号的时空特性以实现有效的神经反应重建。为弥合视觉特征与神经表征之间的模态差异,采用Q-Former将CLIP图像嵌入映射至TSC-VAE隐空间,生成神经代理嵌入。为实现全面的跨模态对齐,结合均方误差(MSE)损失进行逐点特征匹配,以及切片沃瑟斯坦距离(SWD)对神经代理嵌入与TSC-VAE隐空间嵌入的概率分布进行对齐。我们在THINGS-EEG2和THINGS-MEG数据集上进行了大量实验,证明了该方法从视觉刺激生成高质量M/EEG信号的有效性。

原文摘要 · Abstract (English)

Brain encoding models not only serve to decipher how visual stimuli are transformed into neural responses, but also represent a critical step toward visual prostheses that restore vision for patients with severe vision disorders. Brain encoding involves two fundamental steps: achieving faithful reconstruction of neural responses and establishing cross-modal alignment between visual stimuli and neural responses. To this end, we propose ViBE, a novel brain encoding framework for generating magnetoencephalography (MEG) and electroencephalography (EEG) signals from visual stimuli. Specifically, we first design a spatio-temporal convolutional variational autoencoder (TSC-VAE) that captures the spatio-temporal characteristics of M/EEG signals for effective neural response reconstruction. To bridge the modality gap between visual features and neural representations, we employ Q-Former to map CLIP image embeddings to the TSC-VAE latent space, producing neural proxy embeddings. For comprehensive cross-modal alignment, we combine mean squared error (MSE) loss for point-wise feature matching with sliced Wasserstein distance (SWD) for probability distribution alignment between the neural proxy embeddings and TSC-VAE latent embeddings. We conduct extensive experiments on the THINGS-EEG2 and THINGS-MEG datasets, demonstrating the effectiveness of our approach in generating high-quality M/EEG signals from visual stimuli.

脑机接口多模态生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。