arXiv:2602.22555cs.LGcs.AI2026-02被引 2

用自回归方法从脑电波高效还原图像,参数少、效果好。

Autoregressive Visual Decoding from EEG Signals

  • 基于多尺度预测的自回归生成,从脑电到图像逐步细化。
  • 在两个数据集上超越现有方法,参数仅用10%。
  • 生成过程反映人类视觉层级特性,适合实际脑机接口应用。

脑电图(EEG)因成本低、时间分辨率高,成为视觉信息解码的热门媒介。然而,当前方法在脑电与图像之间存在显著模态鸿沟,通常依赖多阶段复杂适配,难以保持一致性且易累积误差。此外,大规模扩散模型带来的计算开销限制了其在真实脑机接口(BCI)中的实用性。本文提出AVDE,一种轻量高效的脑电视觉解码框架。首先,利用预训练的LaBraM模型,通过对比学习微调以对齐脑电与图像表示;其次,采用基于“下一尺度预测”的自回归生成框架:使用预训练VQ-VAE将图像编码为多尺度标记图,再训练一个Transformer从脑电嵌入作为最粗粒度表示出发,自回归预测更精细尺度的标记。该设计实现连贯生成的同时,保持脑电信号与重建图像间的直接关联。在两个数据集上的实验表明,AVDE在图像检索和重建任务中均优于现有最先进方法,且仅使用10%的参数量。此外,中间输出可视化显示,AVDE的生成过程反映了人类视觉感知的层级特性。这些结果凸显自回归模型作为高效、可解释工具在实用脑机接口中的潜力。

原文摘要 · Abstract (English)

Electroencephalogram (EEG) signals have become a popular medium for decoding visual information due to their cost-effectiveness and high temporal resolution. However, current approaches face significant challenges in bridging the modality gap between EEG and image data. These methods typically rely on complex adaptation processes involving multiple stages, making it hard to maintain consistency and manage compounding errors. Furthermore, the computational overhead imposed by large-scale diffusion models limit their practicality in real-world brain-computer interface (BCI) applications. In this work, we present AVDE, a lightweight and efficient framework for visual decoding from EEG signals. First, we leverage LaBraM, a pre-trained EEG model, and fine-tune it via contrastive learning to align EEG and image representations. Second, we adopt an autoregressive generative framework based on a "next-scale prediction" strategy: images are encoded into multi-scale token maps using a pre-trained VQ-VAE, and a transformer is trained to autoregressively predict finer-scale tokens starting from EEG embeddings as the coarsest representation. This design enables coherent generation while preserving a direct connection between the input EEG signals and the reconstructed images. Experiments on two datasets show that AVDE outperforms previous state-of-the-art methods in both image retrieval and reconstruction tasks, while using only 10% of the parameters. In addition, visualization of intermediate outputs shows that the generative process of AVDE reflects the hierarchical nature of human visual perception. These results highlight the potential of autoregressive models as efficient and interpretable tools for practical BCI applications.

脑机接口自回归生成视觉解码EEG

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。