arXiv:2604.10617eess.IVcs.CV2026-04

用脑信号生成图像时,让模型更懂物体位置和结构。

Brain-Grasp: Graph-based Saliency Priors for Improved fMRI-based Visual Brain Decoding

  • 用图结构提取脑信号中的显著性线索,生成空间掩码
  • 结合语义信息,单次引导扩散模型重建图像
  • 提升物体一致性与场景自然度,适合脑解码研究者

基于脑电信号的图像重建近年取得进展,但在保持物体层级结构和语义准确性方面仍面临挑战。现有方法常忽视显著物体的空间布局,导致输出概念不一致。本文提出一种基于图结构的显著性先验解码框架,将脑信号中的结构线索转化为空间掩码,结合嵌入提取的语义信息,条件化一个扩散模型以指导图像再生,从而在保持自然场景构型的同时增强物体一致性。相比多阶段扩散流程,本方法仅依赖单一冻结模型,设计更轻量高效。实验表明,该策略在概念对齐和结构相似性上均有提升,并为高效、可解释且结构化的脑解码提供了新方向。

原文摘要 · Abstract (English)

Recent progress in brain-guided image generation has improved the quality of fMRI-based reconstructions; however, fundamental challenges remain in preserving object-level structure and semantic fidelity. Many existing approaches overlook the spatial arrangement of salient objects, leading to conceptually inconsistent outputs. We propose a saliency-driven decoding framework that employs graph-informed saliency priors to translate structural cues from brain signals into spatial masks. These masks, together with semantic information extracted from embeddings, condition a diffusion model to guide image regeneration, helping preserve object conformity while maintaining natural scene composition. In contrast to pipelines that invoke multiple diffusion stages, our approach relies on a single frozen model, offering a more lightweight yet effective design. Experiments show that this strategy improves both conceptual alignment and structural similarity to the original stimuli, while also introducing a new direction for efficient, interpretable, and structurally grounded brain decoding.

脑解码扩散模型图像重建

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。