用大脑活动重建视频,让画面更连贯、重点更准确。
SemVideo: Reconstructs What You Watch from Brain Activity via Hierarchical Semantic Guidance
- 通过分层语义引导,精准捕捉视觉内容
- 在两个数据集上实现最佳语义对齐与时间一致性
- 适合脑机接口、视觉认知研究者阅读
从脑活动重建动态视觉体验为探索人类视觉感知的神经机制提供了有力途径。尽管基于fMRI的图像重建取得显著进展,但将其拓展至视频重建仍面临重大挑战。现有fMRI到视频的重建方法普遍存在两大缺陷:(i) 关键对象在不同帧间视觉表征不一致,导致外观错位;(ii) 时间连贯性差,出现运动错配或帧间突变。为此,我们提出SemVideo,一种基于分层语义信息引导的新型fMRI到视频重建框架。其核心是SemMiner,一个从原始视频刺激构建三个层次语义线索的模块:静态锚定描述、运动导向叙事和整体摘要。借助该语义引导,SemVideo包含三个关键组件:语义对齐解码器,将fMRI信号与SemMiner生成的CLIP风格嵌入对齐;运动适配解码器,采用创新的三元注意力融合架构重建动态运动模式;条件化视频渲染模块,利用分层语义指导视频重建。在CC2017和HCP数据集上的实验表明,SemVideo在语义对齐与时间一致性方面均表现卓越,创下fMRI到视频重建的新基准。
原文摘要 · Abstract (English)
Reconstructing dynamic visual experiences from brain activity provides a compelling avenue for exploring the neural mechanisms of human visual perception. While recent progress in fMRI-based image reconstruction has been notable, extending this success to video reconstruction remains a significant challenge. Current fMRI-to-video reconstruction approaches consistently encounter two major shortcomings: (i) inconsistent visual representations of salient objects across frames, leading to appearance mismatches; (ii) poor temporal coherence, resulting in motion misalignment or abrupt frame transitions. To address these limitations, we introduce SemVideo, a novel fMRI-to-video reconstruction framework guided by hierarchical semantic information. At the core of SemVideo is SemMiner, a hierarchical guidance module that constructs three levels of semantic cues from the original video stimulus: static anchor descriptions, motion-oriented narratives, and holistic summaries. Leveraging this semantic guidance, SemVideo comprises three key components: a Semantic Alignment Decoder that aligns fMRI signals with CLIP-style embeddings derived from SemMiner, a Motion Adaptation Decoder that reconstructs dynamic motion patterns using a novel tripartite attention fusion architecture, and a Conditional Video Render that leverages hierarchical semantic guidance for video reconstruction. Experiments conducted on the CC2017 and HCP datasets demonstrate that SemVideo achieves superior performance in both semantic alignment and temporal consistency, setting a new state-of-the-art in fMRI-to-video reconstruction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。