arXiv:2603.19667cs.CVcs.AI2026-03

用脑电与文本联合建模,实现更逼真的视觉重建。

Toward High-Fidelity Visual Reconstruction: From EEG-Based Conditioned Generation to Joint-Modal Guided Rebuilding

  • 将脑电与文本视为独立模态,避免信息压缩
  • 多尺度编码+图像增强,提升空间与色彩还原度
  • 适合脑机接口与高保真重建研究者

人类视觉重建旨在基于用户描述和神经信号还原精细视觉刺激。脑电图(EEG)能捕捉丰富的视觉认知信息,包括场景中的复杂空间关系与色彩细节。然而,现有方法依赖对齐框架,强制将EEG特征与文本或图像语义对齐,导致EEG中蕴含的空间与色彩细节被压缩,仅实现条件生成而非高保真重建。为此,本文提出联合模态视觉重建(JMVR)框架,将EEG与文本视为独立模态进行联合学习,以保留EEG特有信息。该框架采用多尺度EEG编码策略,同时捕获细粒度与粗粒度特征,并结合图像增强技术,提升感知细节恢复能力。在THINGS-EEG数据集上的大量实验表明,JMVR优于六种基线方法,尤其在空间结构建模与色彩保真度方面表现突出。

原文摘要 · Abstract (English)

Human visual reconstruction aims to reconstruct fine-grained visual stimuli based on subject-provided descriptions and corresponding neural signals. As a widely adopted modality, Electroencephalography (EEG) captures rich visual cognition information, encompassing complex spatial relationships and chromatic details within scenes. However, current approaches are deeply coupled with an alignment framework that forces EEG features to align with text or image semantic representation. The dependency may condense the rich spatial and chromatic details in EEG that achieved mere conditioned image generation rather than high-fidelity visual reconstruction. To address this limitation, we propose a novel Joint-Modal Visual Reconstruction (JMVR) framework. It treats EEG and text as independent modalities for joint learning to preserve EEG-specific information for reconstruction. It further employs a multi-scale EEG encoding strategy to capture both fine- and coarse-grained features, alongside image augmentation to enhance the recovery of perceptual details. Extensive experiments on the THINGS-EEG dataset demonstrate that JMVR achieves SOTA performance against six baseline methods, specifically exhibiting superior capabilities in modeling spatial structure and chromatic fidelity.

视觉重建脑电图多模态高保真

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。