arXiv:2604.03619cs.CV2026-04

用自然图像自动编码器压缩脑影像,实现高效长时序建模

Can Natural Image Autoencoders Compactly Tokenize fMRI Volumes for Long-Range Dynamics Modeling?

  • 用预训练2D图像自编码器将3D脑影像转为紧凑连续令牌
  • 在多个数据集上超越现有模型,内存占用降低80%以上
  • 适合需要高效建模大脑长程动态的研究者

由于功能性磁共振成像(fMRI)四维信号维度极高,建模其长时序时空动态仍是关键挑战。传统基于体素的模型虽性能优异且可解释性强,但受限于极高的内存需求,仅能捕捉有限时间窗口。为此,我们提出TABLeT(Two-dimensionally Autoencoded Brain Latent Transformer),利用预训练的2D自然图像自编码器对fMRI体积进行编码,将每个3D fMRI体积压缩为一组紧凑的连续令牌,使仅需少量显存的简单Transformer编码器即可实现长序列建模。在英国生物银行(UKB)、人类连接组计划(HCP)和ADHD-200等大规模基准测试中,TABLeT在多项任务上优于现有模型,并在相同输入条件下相较当前最优体素方法显著提升计算与内存效率。此外,我们设计了自监督掩码令牌建模方法对TABLeT进行预训练,进一步提升了其在多种下游任务的表现。结果表明该方法为可扩展、可解释的大脑活动时空建模提供了新路径。

原文摘要 · Abstract (English)

Modeling long-range spatiotemporal dynamics in functional Magnetic Resonance Imaging (fMRI) remains a key challenge due to the high dimensionality of the four-dimensional signals. Prior voxel-based models, although demonstrating excellent performance and interpretation capabilities, are constrained by prohibitive memory demands and thus can only capture limited temporal windows. To address this, we propose TABLeT (Two-dimensionally Autoencoded Brain Latent Transformer), a novel approach that tokenizes fMRI volumes using a pre-trained 2D natural image autoencoder. Each 3D fMRI volume is compressed into a compact set of continuous tokens, enabling long-sequence modeling with a simple Transformer encoder with limited VRAM. Across large-scale benchmarks including the UK-Biobank (UKB), Human Connectome Project (HCP), and ADHD-200 datasets, TABLeT outperforms existing models in multiple tasks, while demonstrating substantial gains in computational and memory efficiency over the state-of-the-art voxel-based method given the same input. Furthermore, we develop a self-supervised masked token modeling approach to pre-train TABLeT, which improves the model's performance for various downstream tasks. Our findings suggest a promising approach for scalable and interpretable spatiotemporal modeling of brain activity. Our code is available at https://github.com/beotborry/TABLeT.

脑影像建模自编码器Transformer长时序分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。