arXiv:2507.02311cs.CV2025-07

用fMRI信号干预图像特征,提升视觉任务精度。

Perception Activator: An intuitive and portable framework for brain cognitive exploration

  • 通过交叉注意力将fMRI特征注入多尺度图像特征
  • 在目标检测与实例分割上准确率显著提升
  • 揭示fMRI蕴含丰富语义与粗略空间信息

脑-视觉解码的最新进展已实现从人类视觉皮层的神经活动(如功能磁共振成像,fMRI)中高保真重建感知的视觉刺激。现有方法多采用像素级与语义级双层策略,但过度依赖低层级像素对齐,缺乏充分精细的语义对齐,导致多个语义对象的重建存在明显失真。为此,我们构建了一个实验框架,以fMRI表征作为干预条件,通过交叉注意力将其注入多尺度图像特征,对比加入与不加入fMRI信息时,在目标检测与实例分割任务上的下游性能及中间特征变化。结果表明,引入fMRI信号可显著提升下游任务的检测与分割精度,证实了fMRI包含丰富的多对象语义线索和粗粒度空间定位信息——这些是当前模型尚未充分挖掘与整合的要素。

原文摘要 · Abstract (English)

Recent advances in brain-vision decoding have driven significant progress, reconstructing with high fidelity perceived visual stimuli from neural activity, e.g., functional magnetic resonance imaging (fMRI), in the human visual cortex. Most existing methods decode the brain signal using a two-level strategy, i.e., pixel-level and semantic-level. However, these methods rely heavily on low-level pixel alignment yet lack sufficient and fine-grained semantic alignment, resulting in obvious reconstruction distortions of multiple semantic objects. To better understand the brain's visual perception patterns and how current decoding models process semantic objects, we have developed an experimental framework that uses fMRI representations as intervention conditions. By injecting these representations into multi-scale image features via cross-attention, we compare both downstream performance and intermediate feature changes on object detection and instance segmentation tasks with and without fMRI information. Our results demonstrate that incorporating fMRI signals enhances the accuracy of downstream detection and segmentation, confirming that fMRI contains rich multi-object semantic cues and coarse spatial localization information-elements that current models have yet to fully exploit or integrate.

脑机接口视觉解码fMRI多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。