arXiv:2511.02565cs.CVcs.AI2025-11中稿 · ICLR被引 3

模仿大脑视觉系统结构,实现无需训练即可还原他人视觉的快速解码

A Cognitive Process-Inspired Architecture for Subject-Agnostic Brain Visual Decoding

  • 基于人脑视觉通路分层设计解码架构,分离并利用早期视觉与背侧/腹侧流特征
  • 在未见受试者上实现93%重建准确率,单次生成仅需10秒,无需重新训练
  • 适合临床脑机接口、无感视觉重建等需快速泛化应用的场景

无受试者特异性脑解码旨在不依赖个体训练的情况下,从fMRI数据中重建连续视觉体验,具有重要的临床潜力。然而,跨受试者泛化困难和脑信号复杂性使得该方向仍处于探索阶段。本文提出视觉皮层流架构(VCFlow),一种显式建模人类视觉系统背侧-腹侧结构的分层解码框架,用于学习多维表征。通过解耦并利用来自初级视觉皮层、腹侧流和背侧流的特征,VCFlow捕获了对视觉重建至关重要的多样化互补认知信息。此外,我们引入特征级对比学习策略,增强对受试者无关语义表征的提取,从而提升对未见受试者的泛化能力。与传统方法需每受试者超过12小时数据及大量计算不同,VCFlow平均仅损失7%准确率,但可实现每帧视频10秒内生成且无需重训练,提供了一种快速且临床可扩展的解决方案。代码将在论文接收后公开。

原文摘要 · Abstract (English)

Subject-agnostic brain decoding, which aims to reconstruct continuous visual experiences from fMRI without subject-specific training, holds great potential for clinical applications. However, this direction remains underexplored due to challenges in cross-subject generalization and the complex nature of brain signals. In this work, we propose Visual Cortex Flow Architecture (VCFlow), a novel hierarchical decoding framework that explicitly models the ventral-dorsal architecture of the human visual system to learn multi-dimensional representations. By disentangling and leveraging features from early visual cortex, ventral, and dorsal streams, VCFlow captures diverse and complementary cognitive information essential for visual reconstruction. Furthermore, we introduce a feature-level contrastive learning strategy to enhance the extraction of subject-invariant semantic representations, thereby enhancing subject-agnostic applicability to previously unseen subjects. Unlike conventional pipelines that need more than 12 hours of per-subject data and heavy computation, VCFlow sacrifices only 7\% accuracy on average yet generates each reconstructed video in 10 seconds without any retraining, offering a fast and clinically scalable solution. The source code will be released upon acceptance of the paper.

脑解码视觉重建无监督fMRI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。