arXiv:2605.01240cs.LGcs.AI2026-05

Rhamba通过区域感知掩码与混合架构,提升静息态脑影像自监督学习效果。

Rhamba: Region-Aware Hybrid Attention-Mamba Framework for Self-Supervised Learning in Resting-State fMRI

论文配图:Rhamba: Region-Aware Hybrid Attention-Mamba Framework for Self-Supervised Learning in Resting-State fMRI
图 1 · 摘自论文原文
  • 引入解剖引导的区域感知掩码,结合注意力与Mamba混合结构建模脑区序列。
  • MA混合架构在双数据集上平均AUROC达最优,显著优于现有方法。
  • 揭示掩码策略与模型结构交互影响性能,提升可解释性与泛化能力。

自监督预训练在大规模神经影像中前景广阔,但区域感知掩码与混合序列建模的影响尚未充分探索。本文提出Rhamba框架,结合解剖学引导的掩码策略与混合注意力-Mamba架构,用于静息态功能磁共振成像(fMRI)分析。模型在ABIDE数据集上使用区域对齐的块嵌入与三种掩码策略(任意、多数、纯)进行预训练,涵盖从低到高的空间特异性。评估四种架构变体:仅Mamba、交替式(交替嵌套Mamba与注意力块)、两种混合编码器-解码器结构(注意力-Mamba, AM;Mamba-Attention, MA)。预训练模型在COBRE与ADHD-200数据集上微调,用于精神分裂症与多动症分类任务。采用集成梯度法识别关键脑区。掩码策略显著影响重建行为,重建损失呈一致趋势(任意 > 多数 > 纯),但下游性能差异较小且依赖数据集。其中MA配置在两数据集上平均AUROC最高,整体表现超越当前最优方法。区域分析表明,最佳性能取决于掩码策略与架构的协同作用,而非单一最优组合。总体而言,Rhamba为大规模fMRI表征学习提供了兼顾可解释性、可扩展性与性能的灵活框架。

原文摘要 · Abstract (English)

Self-supervised pretraining is promising for large-scale neuroimaging, yet the impact of region-aware masking and hybrid sequence modeling remains underexplored. In this work, we introduce Rhamba, a region-aware pretraining framework that integrates anatomically guided masking with hybrid Attention-Mamba architectures for resting state functional magnetic resonance imaging (fMRI) analysis. Models were pretrained on the ABIDE dataset using region-aligned patch embeddings and three masking strategies (Any, Majority, and Pure) with increasing spatial specificity. We evaluated four architectural variants: a Mamba only model, an Alternate architecture with interleaved Mamba and Attention blocks, and two hybrid encoder-decoder configurations (Attention-Mamba (AM) and Mamba-Attention (MA)). The pretrained models were fine-tuned on downstream classification tasks using the COBRE and ADHD-200 datasets for schizophrenia and attention-deficit/hyperactivity disorder discrimination. We employed Integrated Gradients, an explainable AI method, to identify the brain regions contributing to model predictions. Masking strategy strongly influenced reconstruction behavior, with reconstruction loss following a consistent ordering (Any > Majority > Pure). However, this trend did not directly translate into downstream performance, where differences were modest and dataset-dependent. The hybrid architecture with the MA configuration achieved the highest average AUROC across both datasets, and Rhamba outperformed state-of-the-art methods in comparative evaluation. Region-wise analysis showed that peak performance depends on the interaction between masking strategy and architecture rather than a single dominant configuration. Overall, Rhamba offers a flexible framework for balancing interpretability, scalability, and performance in large-scale fMRI representation learning.

fMRI自监督混合架构可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。