用细粒度文本作桥梁,提升脑电到图像的重建细节与语义一致性。
Improving Brain-to-Image Reconstruction via Fine-Grained Text Bridging
- 用视觉语言大模型生成刺激图像的细粒度描述,作为脑活动与图像间的桥梁。
- 提出三项奖励指标,引导模型从fMRI信号解码出准确的细粒度文本描述。
- 可无缝集成至现有方法,适合关注脑机接口语义重建的研究者。
脑到图像重建旨在从大脑活动恢复人类感知的视觉刺激。然而,重建结果常缺乏细节且存在语义不一致,可能源于语义信息不足。为此,我们提出细粒度脑到图像重建方法(FgB2I),通过细粒度文本作为桥梁提升重建质量。FgB2I包含三个阶段:细节增强、细粒度文本描述解码、文本桥接的脑到图像重建。在细节增强阶段,利用大规模视觉语言模型生成视觉刺激的细粒度标题,并实验验证其有效性。我们设计三种奖励指标(物体准确性、文本-图像语义相似性、图像-图像语义相似性)以指导语言模型从fMRI信号中解码细粒度文本描述。这些细粒度文本可被整合进现有重建方法,实现更精细的脑到图像重建。
原文摘要 · Abstract (English)
Brain-to-Image reconstruction aims to recover visual stimuli perceived by humans from brain activity. However, the reconstructed visual stimuli often missing details and semantic inconsistencies, which may be attributed to insufficient semantic information. To address this issue, we propose an approach named Fine-grained Brain-to-Image reconstruction (FgB2I), which employs fine-grained text as bridge to improve image reconstruction. FgB2I comprises three key stages: detail enhancement, decoding fine-grained text descriptions, and text-bridged brain-to-image reconstruction. In the detail-enhancement stage, we leverage large vision-language models to generate fine-grained captions for visual stimuli and experimentally validate its importance. We propose three reward metrics (object accuracy, text-image semantic similarity, and image-image semantic similarity) to guide the language model in decoding fine-grained text descriptions from fMRI signals. The fine-grained text descriptions can be integrated into existing reconstruction methods to achieve fine-grained Brain-to-Image reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。