用双向编码提升跨人脑图像解码精度,解决误差累积问题
Cross-Subject Mind Decoding from Inaccurate Representations
- 双向映射+双向自编码器,统一多受试者认知差异
- 在多个数据集上优于现有方法,新受试者仅需少量样本即可适配
- 引入语义优化与视觉一致性模块,提升重建图像质量
从fMRI信号中解码刺激图像已借助预训练生成模型取得进展。然而,由于认知差异和个体特异性,现有方法在跨受试者映射上表现不佳。问题源于序列误差:单向映射生成部分不准确的表示,输入扩散模型后误差累积,导致重建质量下降。为此,我们提出双向自编码器交织框架(Bidirectional Autoencoder Intertwining),通过受试者偏置调制模块统一多受试者特征,并利用双向映射更准确捕捉数据分布。为进一步提升表示解码为图像时的保真度,引入语义精炼模块增强语义表征,以及视觉一致性模块缓解视觉表征不准确的影响。该方法与ControlNet和Stable Diffusion结合,在基准数据集上实现定性和定量双优表现。此外,框架对新受试者具有强适应性,仅需极少训练样本即可有效运行。
原文摘要 · Abstract (English)
Decoding stimulus images from fMRI signals has advanced with pre-trained generative models. However, existing methods struggle with cross-subject mappings due to cognitive variability and subject-specific differences. This challenge arises from sequential errors, where unidirectional mappings generate partially inaccurate representations that, when fed into diffusion models, accumulate errors and degrade reconstruction fidelity. To address this, we propose the Bidirectional Autoencoder Intertwining framework for accurate decoded representation prediction. Our approach unifies multiple subjects through a Subject Bias Modulation Module while leveraging bidirectional mapping to better capture data distributions for precise representation prediction. To further enhance fidelity when decoding representations into stimulus images, we introduce a Semantic Refinement Module to improve semantic representations and a Visual Coherence Module to mitigate the effects of inaccurate visual representations. Integrated with ControlNet and Stable Diffusion, our method outperforms state-of-the-art approaches on benchmark datasets in both qualitative and quantitative evaluations. Moreover, our framework exhibits strong adaptability to new subjects with minimal training samples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。