无需训练即可还原他人脑活动成图像,突破个体差异限制。
The Pictorial Cortex: Zero-Shot Cross-Subject fMRI-to-Image Reconstruction via Compositional Latent Modeling
- 用组合潜空间建模脑活动,分解个体/实验/刺激差异
- 零样本跨被试重建效果优于现有方法,生成图像更清晰
- 适合神经科学与脑机接口研究者,推动通用视觉解码
从人类脑活动解码视觉体验是神经科学、神经影像与人工智能交叉的核心挑战。关键难点在于皮层反应的固有变异性:相同视觉刺激在不同个体和试验中引发的神经活动存在差异,导致fMRI到图像的重建非单射。本文解决一个极具实际意义但极具挑战的问题——零样本跨被试fMRI到图像重建,即对未见过的被试进行无特定训练的视觉体验重建。为实现严谨评估,我们构建了统一皮层表面数据集UniCortex-fMRI,整合多个视觉刺激fMRI数据集,覆盖广泛被试与刺激。该数据集经标准化处理,支持零样本跨被试重建的探索。针对建模挑战,我们提出PictorialCortex,采用组合潜变量形式化建模,将受试者、数据集和试验相关变异性下的刺激驱动表征结构化。PictorialCortex在通用皮层潜空间中运行,通过潜因子分解-组合模块实现该形式化,并引入配对分解与重分解一致性正则化。推理时,聚合多个已见被试条件下的代理潜变量,指导扩散模型生成未知被试的图像。大量实验表明,PictorialCortex显著提升零样本跨被试视觉重建性能,凸显组合潜变量建模与多数据集训练的优势。
原文摘要 · Abstract (English)
Decoding visual experiences from human brain activity remains a central challenge at the intersection of neuroscience, neuroimaging, and artificial intelligence. A critical obstacle is the inherent variability of cortical responses: neural activity elicited by the same visual stimulus differs across individuals and trials due to anatomical, functional, cognitive, and experimental factors, making fMRI-to-image reconstruction non-injective. In this paper, we tackle a challenging yet practically meaningful problem: zero-shot cross-subject fMRI-to-image reconstruction, where the visual experience of a previously unseen individual must be reconstructed without subject-specific training. To enable principled evaluation, we present a unified cortical-surface dataset -- UniCortex-fMRI, assembled from multiple visual-stimulus fMRI datasets to provide broad coverage of subjects and stimuli. Our UniCortex-fMRI is particularly processed by standardized data formats to make it possible to explore this possibility in the zero-shot scenario of cross-subject fMRI-to-image reconstruction. To tackle the modeling challenge, we propose PictorialCortex, which models fMRI activity using a compositional latent formulation that structures stimulus-driven representations under subject-, dataset-, and trial-related variability. PictorialCortex operates in a universal cortical latent space and implements this formulation through a latent factorization-composition module, reinforced by paired factorization and re-factorizing consistency regularization. During inference, surrogate latents synthesized under multiple seen-subject conditions are aggregated to guide diffusion-based image synthesis for unseen subjects. Extensive experiments show that PictorialCortex improves zero-shot cross-subject visual reconstruction, highlighting the benefits of compositional latent modeling and multi-dataset training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。