arXiv:2604.08068cs.CV2026-04

用脑电波生成3D视觉模型,突破传统2D重建限制

Brain3D: EEG-to-3D Decoding of Visual Representations via Multimodal Reasoning

  • 分步构建:先从脑电生成图像,再用大模型提取3D描述
  • 3D重建达到85.4%解码准确率,CLIPScore达0.648
  • 适合脑机接口、虚拟现实等需3D建模的场景

从脑电图(EEG)解码视觉信息近年来取得显著进展,主要集中在从脑活动重建二维(2D)图像。然而,三维(3D)表示的重建仍基本未被探索,这限制了神经解码在几何理解上的应用。为此,我们提出Brain3D,一种基于多模态推理的EEG到3D重建架构,其以EEG到图像解码为基础。该方法通过几何感知的生成推理,逐步将神经表征转换至3D域。流程首先从EEG信号生成视觉上一致的图像,随后利用多模态大语言模型提取结构化的3D感知描述,引导基于扩散模型的生成阶段,最终通过单图像转3D模型将输出转换为连贯的3D网格。通过分阶段设计,该方法避免直接的EEG到3D映射,实现可扩展的脑驱动3D生成。我们在多个原图与重建结果间进行综合评估,考察语义对齐与几何保真度。实验表明该架构表现优异,最高实现85.4%的10分类别Top-1 EEG解码准确率和0.648的CLIPScore,验证了多模态脑驱动3D重建的可行性。

原文摘要 · Abstract (English)

Decoding visual information from electroencephalography (EEG) has recently achieved promising results, primarily focusing on reconstructing two-dimensional (2D) images from brain activity. However, the reconstruction of three-dimensional (3D) representations remains largely unexplored. This limits the geometric understanding and reduces the applicability of neural decoding in different contexts. To address this gap, we propose Brain3D, a multimodal architecture for EEG-to-3D reconstruction based on EEG-to-image decoding. It progressively transforms neural representations into the 3D domain using geometry-aware generative reasoning. Our pipeline first produces visually grounded images from EEG signals, then employs a multimodal large language model to extract structured 3D-aware descriptions, which guide a diffusion-based generation stage whose outputs are finally converted into coherent 3D meshes via a single-image-to-3D model. By decomposing the problem into structured stages, the proposed approach avoids direct EEG-to-3D mappings and enables scalable brain-driven 3D generation. We conduct a comprehensive evaluation comparing the reconstructed 3D outputs against the original visual stimuli, assessing both semantic alignment and geometric fidelity. Experimental results demonstrate strong performance of the proposed architecture, achieving up to 85.4% 10-way Top-1 EEG decoding accuracy and 0.648 CLIPScore, supporting the feasibility of multimodal EEG-driven 3D reconstruction.

脑机接口3D生成多模态神经解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。