arXiv:2601.21673cs.CV2026-01被引 1

用2D模型高效提取3D脑影像特征,提升阿尔茨海默病分类准确率

Multimodal Visual Surrogate Compression for Alzheimer's Disease Classification

  • 将3D脑影像压缩为2D视觉代理特征,结合文本引导捕捉跨切片全局信息
  • 在三个大规模数据集上实现优于现有方法的二分类与多分类性能
  • 适合需要高效处理高维医学影像的临床辅助诊断研究者

高维结构磁共振成像(sMRI)广泛用于阿尔茨海默病(AD)诊断。现有方法多依赖3D架构(如3D CNN)、切片级特征提取后聚合,或使用2D基础模型(如DINO)进行无训练特征提取,但分别存在计算成本高、丢失跨切片关系、判别性特征提取能力弱的问题。为此,我们提出多模态视觉代理压缩(MVSC),将大尺寸3D sMRI体积压缩为紧凑的2D特征(称为视觉代理),使其更适配冻结的2D基础模型以提取强表示,用于最终的AD分类。MVSC包含两个关键组件:基于文本引导的体积上下文编码器,用于捕获全局跨切片上下文;以及文本增强的、分块级的自适应切片融合模块,用于聚合切片级信息。在三个大规模阿尔茨海默病基准数据集上的大量实验表明,相较于现有最优方法,我们的MVSC在二分类和多分类任务中均表现更优。

原文摘要 · Abstract (English)

High-dimensional structural MRI (sMRI) images are widely used for Alzheimer's Disease (AD) diagnosis. Most existing methods for sMRI representation learning rely on 3D architectures (e.g., 3D CNNs), slice-wise feature extraction with late aggregation, or apply training-free feature extractions using 2D foundation models (e.g., DINO). However, these three paradigms suffer from high computational cost, loss of cross-slice relations, and limited ability to extract discriminative features, respectively. To address these challenges, we propose Multimodal Visual Surrogate Compression (MVSC). It learns to compress and adapt large 3D sMRI volumes into compact 2D features, termed as visual surrogates, which are better aligned with frozen 2D foundation models to extract powerful representations for final AD classification. MVSC has two key components: a Volume Context Encoder that captures global cross-slice context under textual guidance, and an Adaptive Slice Fusion module that aggregates slice-level information in a text-enhanced, patch-wise manner. Extensive experiments on three large-scale Alzheimer's disease benchmarks demonstrate our MVSC performs favourably on both binary and multi-class classification tasks compared against state-of-the-art methods.

医学影像多模态压缩阿尔茨海默病

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。