arXiv:2605.23137eess.IVcs.CV2026-05被引 1

通过分阶段对齐,提升脑电到视觉的解码精度与稳定性。

STAMBRIDGE: Spectral-Temporal Amplitude-aware Mid-Feature Bridge for EEG Visual Decoding

论文配图:STAMBRIDGE: Spectral-Temporal Amplitude-aware Mid-Feature Bridge for EEG Visual Decoding
图 1 · 摘自论文原文
  • 先用频时幅感知模块增强脑电信号质量,减少失真。
  • 在中间层构建正则化语义桥,实现稳定跨模态对齐。
  • 零样本检索达34.5%准确率,可生成语义连贯图像。

脑电图(EEG)视觉解码因神经信号信噪比低与视觉-语言空间高度结构化之间的模态差距而面临挑战,导致直接跨模态对齐不稳定。为此,我们提出STAMBRIDGE,一种两阶段通用框架,依次解决特征调节与跨模态对齐问题。首先,引入频时幅感知调制(STAM),通过幅度驱动的软通道加权和多尺度时间卷积替代硬频率掩码,显式保留频率感知瞬态,同时降低时域振铃伪影风险。基于这些稳健的神经特征,进一步提出模型无关的中层语义桥(MFSB),通过定向跨模态交互构建正则化中间空间,实现分阶段知识蒸馏与更稳定的语义对齐。在THINGS-EEG基准上的实验显示,其零样本200类检索性能达到34.50% Top-1与65.95% Top-5准确率。此外,所学嵌入通过扩散模型生成语义一致的图像重建,验证了脑电到视觉的鲁棒语义对齐能力。代码已开源:https://github.com/thabeatmjh/STAMBRIDGE。

原文摘要 · Abstract (English)

Electroencephalography (EEG) visual decoding remains challenging due to the modality gap between low-SNR neural signals and highly structured vision--language spaces, making direct cross-modal alignment unstable. To address this, we propose STAMBRIDGE, a versatile two-stage framework that sequentially tackles feature conditioning and cross-modal alignment. First, we introduce a Spectral-Temporal Amplitude-aware Modulation (STAM) to extract well-conditioned EEG representations. By replacing hard frequency masking with amplitude-derived soft channel weighting and multi-scale temporal convolutions, STAM explicitly preserves frequency-aware transients while reducing the risk of time-domain ringing artifacts. Building upon these robust neural features, we further introduce a model-agnostic Mid-Feature Semantic Bridge (MFSB) that constructs a regularized intermediate space through directed cross-modal interactions, enabling staged distillation and more stable semantic alignment. Experiments on the THINGS-EEG benchmark show competitive 200-way zero-shot retrieval performance, with 34.50\% Top-1 and 65.95\% Top-5 accuracy. In addition, embeddings learned by STAMBRIDGE produce semantically coherent image reconstructions with a diffusion model, demonstrating robust EEG-to-vision semantic alignment. The code is available at: https://github.com/thabeatmjh/STAMBRIDGE.

脑电解码跨模态对齐扩散模型神经信号

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。