arXiv:2504.00432cs.CV2025-04被引 3

将视觉信息分解为语义、空间、运动三部分,更真实还原大脑处理视频的方式。

DecoFuse: Decomposing and Fusing the "What", "Where", and "How" for Brain-Inspired fMRI-to-Video Decoding

  • 分三步解码:先拆解视频为语义、空间、运动三成分,再分别重建并融合。
  • 语义分类准确率82.4%,空间一致性70.6%,运动预测相似度0.212,视频生成50类准确率21.9%。
  • 符合大脑背侧与腹侧通路分工,适合脑科学与生成模型交叉研究者。

从脑活动解码视觉体验是一项重大挑战。现有fMRI-to-video方法多关注语义内容,忽视空间和运动信息,而这些在大脑中通过不同路径处理。受此启发,我们提出DecoFuse,一种受大脑启发的新型fMRI-to-video解码框架。该方法将视频分解为语义、空间和运动三个成分,分别解码后融合重建。这一策略不仅将复杂任务拆分为可管理子任务,还增强了学习表征与生物对应物之间的联系,实验验证其优于现有最优方法:语义分类准确率达82.4%,空间一致性达70.6%,运动预测余弦相似度为0.212,视频生成50类准确率为21.9%。神经编码分析显示语义与空间信息分别对应两流假说,进一步验证了背侧与腹侧通路的独立作用。总体而言,DecoFuse为fMRI-to-video解码提供了强生物学依据的框架。

原文摘要 · Abstract (English)

Decoding visual experiences from brain activity is a significant challenge. Existing fMRI-to-video methods often focus on semantic content while overlooking spatial and motion information. However, these aspects are all essential and are processed through distinct pathways in the brain. Motivated by this, we propose DecoFuse, a novel brain-inspired framework for decoding videos from fMRI signals. It first decomposes the video into three components - semantic, spatial, and motion - then decodes each component separately before fusing them to reconstruct the video. This approach not only simplifies the complex task of video decoding by decomposing it into manageable sub-tasks, but also establishes a clearer connection between learned representations and their biological counterpart, as supported by ablation studies. Further, our experiments show significant improvements over previous state-of-the-art methods, achieving 82.4% accuracy for semantic classification, 70.6% accuracy in spatial consistency, a 0.212 cosine similarity for motion prediction, and 21.9% 50-way accuracy for video generation. Additionally, neural encoding analyses for semantic and spatial information align with the two-streams hypothesis, further validating the distinct roles of the ventral and dorsal pathways. Overall, DecoFuse provides a strong and biologically plausible framework for fMRI-to-video decoding. Project page: https://chongjg.github.io/DecoFuse/.

脑机接口视频生成fMRI解码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。