用大脑双通路机制,让脑扫描图生成更懂语义的视频。
Bridging Brain and Semantics: A Hierarchical Framework for Semantically Enhanced fMRI-to-Video Reconstruction

- 分两步走:先从脑信号提取丰富语义,再调用记忆融合优化
- 在两个数据集上超越现有方法,视频重建更准确
- 适合做神经影像与视频生成交叉研究的人看
从功能性磁共振成像(fMRI)重建动态视觉体验为视频,对理解神经过程至关重要。然而,当前fMRI-to-video重建方法因脑信号与视频内容间存在语义鸿沟而受限,主要源于不完整的语义嵌入,既无法捕捉视频特有线索(如动作),也缺乏先验知识整合。为此,我们受人类大脑双通路处理机制启发,提出CineNeuron,一种用于语义增强的fMRI-to-video重建的层次化框架,包含两个协同阶段:首先,自下而上的语义增强阶段将fMRI信号映射到一个全面捕捉文本语义、图像内容、动作概念和物体类别的丰富嵌入空间;其次,自上而下的记忆集成阶段利用提出的混合记忆(Mixture-of-Memories)方法,动态选择先前见过的数据中的相关“记忆”,并将其与fMRI嵌入融合以精炼视频重建。在两个fMRI-to-video基准上的大量实验结果表明,CineNeuron在多种指标上均优于现有最先进方法。
原文摘要 · Abstract (English)
Reconstructing dynamic visual experiences as videos from functional magnetic resonance imaging (fMRI) is pivotal for advancing the understanding of neural processes. However, current fMRI-to-video reconstruction methods are hindered by a semantic gap between noisy fMRI signals and the rich content of videos, stemming from a reliance on incomplete semantic embeddings that neither capture video-specific cues (e.g., actions) nor integrate prior knowledge. To this end, we draw inspiration from the dual-pathway processing mechanism in human brain and introduce CineNeuron, a novel hierarchical framework for semantically enhanced video reconstruction from fMRI signals with two synergistic stages. First, a bottom-up semantic enrichment stage maps fMRI signals to a rich embedding space that comprehensively captures textual semantics, image contents, action concepts, and object categories. Second, a top-down memory integration stage utilizes the proposed Mixture-of-Memories method to dynamically select relevant "memories" from previously seen data and fuse them with the fMRI embedding to refine the video reconstruction. Extensive experimental results on two fMRI-to-video benchmarks demonstrate that CineNeuron surpasses state-of-the-art methods across various metrics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。