arXiv:2601.17857cs.CV2026-01被引 2

用语义描述提升脑图图像重建准确性,减少幻觉。

SynMind: Reducing Semantic Hallucination in fMRI-Based Image Reconstruction

  • 将脑信号解析为句子级语义描述,增强对象识别
  • 在多个指标上优于现有方法,用小模型单卡实现
  • 适合关注图像重建真实性和可解释性的研究者

基于fMRI的图像重建虽已实现高度逼真的视觉效果,但常出现语义错位——关键物体被替换或幻化。本文提出SynMind框架,重新思考显式语义理解在解码中的作用。现有方法过度依赖低层次视觉特征(如纹理与整体构图),而忽略对象身份。为此,利用接地视觉语言模型生成多层次、类人化的文本表征,捕捉物体身份与空间关系。SynMind将这些显式语义编码与视觉先验结合,驱动预训练扩散模型重建图像。大量实验表明,该方法在多数定量指标上超越当前最优方案。尤为突出的是,仅用Stable Diffusion 1.4和单张消费级显卡,即优于基于SDXL的方法。大规模人工评估显示,重建结果更符合人类视觉感知。神经可视化分析发现,该方法激活更广泛且语义相关的脑区,缓解对高层视觉区域的过度依赖。

原文摘要 · Abstract (English)

Recent advances in fMRI-based image reconstruction have achieved remarkable photo-realistic fidelity. Yet, a persistent limitation remains: while reconstructed images often appear naturalistic and holistically similar to the target stimuli, they frequently suffer from severe semantic misalignment -- salient objects are often replaced or hallucinated despite high visual quality. In this work, we address this limitation by rethinking the role of explicit semantic interpretation in fMRI decoding. We argue that existing methods rely too heavily on entangled visual embeddings which prioritize low-level appearance cues -- such as texture and global gist -- over explicit semantic identity. To overcome this, we parse fMRI signals into rich, sentence-level semantic descriptions that mirror the hierarchical and compositional nature of human visual understanding. We achieve this by leveraging grounded VLMs to generate synthetic, human-like, multi-granularity textual representations that capture object identities and spatial organization. Built upon this foundation, we propose SynMind, a framework that integrates these explicit semantic encodings with visual priors to condition a pretrained diffusion model. Extensive experiments demonstrate that SynMind outperforms state-of-the-art methods across most quantitative metrics. Notably, by offloading semantic reasoning to our text-alignment module, SynMind surpasses competing methods based on SDXL while using the much smaller Stable Diffusion 1.4 and a single consumer GPU. Large-scale human evaluations further confirm that SynMind produces reconstructions more consistent with human visual perception. Neurovisualization analyses reveal that SynMind engages broader and more semantically relevant brain regions, mitigating the over-reliance on high-level visual areas.

fMRI重建语义对齐扩散模型脑机接口

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。