arXiv:2605.18172cs.AI2026-05

用生成图像还原脑电隐含视觉信息,提升多模态大模型理解能力

Visualizing the Invisible: Generative Visual Grounding Empowers Universal EEG Understanding in MLLMs

论文配图:Visualizing the Invisible: Generative Visual Grounding Empowers Universal EEG Understanding in MLLMs
图 1 · 摘自论文原文
  • 通过脑电到图像生成模型,将非视觉脑电信号转化为视觉代理图像
  • 轻量级模型仅调1.7亿参数即达70亿参数文本对齐基线效果
  • 适合医疗脑状态分析、脑机接口与多模态认知建模研究者

预训练语言模型和多模态大模型(MLLMs)为脑基础模型提供了可能,但视觉诱发的脑电数据集仍稀缺,现有方法主要将神经信号对齐抽象文本,损失了脑活动中的精细感知信息。本文提出生成式视觉定位(GVG)框架,利用脑电到图像生成模型作为视觉翻译器,将非视觉脑电信号转换为实例相关的视觉代理图像,使MLLMs能借助其视觉先验进行临床状态解析。在两个MLLM骨干网络(GVG-X-Omni与GVG-Janus)上验证,仅图像对齐已具竞争力:轻量级的GVG-X-Omni仅微调1.7亿参数,即达到调优70亿参数文本对齐基线的效果。进一步扩展的三模态(图像+文本)对齐中,文本提供语义锚点,视觉代理增强神经表征的感知细节。实验显示在脑电理解与视觉生成上均有稳定提升,表明视觉代理对齐是文本对齐的有效补充。

原文摘要 · Abstract (English)

Leveraging the universal representations of pre-trained LLMs and MLLMs offers a promising path toward brain foundation models. However, visually-evoked EEG datasets remain scarce, leading existing methods to align neural signals mainly with abstract text, a lossy translation that may discard fine-grained perceptual information encoded in brain activity. We propose Generative Visual Grounding (GVG), a framework that visualizes the invisible by using an EEG-to-image generative model as a visual translator. Instead of forcing EEG into text alone, GVG hallucinates instance-specific proxy images for non-visual EEG, providing structured visual contexts that allow MLLMs to exploit their visual priors for clinical-state interpretation. We validate this idea on two MLLM backbones, GVG-X-Omni and GVG-Janus. Image-only alignment is already competitive: the lightweight GVG-X-Omni matches 1.7B-parameter text-aligned baselines while tuning only 170M parameters on a frozen 7B backbone. We further extend GVG-Janus with trimodal Image+Text alignment, where text supplies categorical semantic anchors and visual proxies enrich neural representations with perceptual details. Experiments show consistent gains in EEG understanding and visual generation, suggesting visual proxy grounding as an effective complement to textual alignment.

脑电分析多模态生成模型视觉定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。