arXiv:2605.16775cs.CVcs.AI2026-05中稿 · EMBC 2026被引 1

Volta-3D通过联合对齐全局与局部特征,提升脑部MRI模型的跨任务泛化能力。

VolTA-3D: Self-Supervised Learning for Brain MRI using 3D Volumetric Token Alignment

论文配图:VolTA-3D: Self-Supervised Learning for Brain MRI using 3D Volumetric Token Alignment
图 1 · 摘自论文原文
  • 采用师生架构联合对齐全局风格与局部补丁特征。
  • 在阿尔茨海默病分类等任务上超越随机初始化基线。
  • 适合需要多任务迁移的临床影像分析场景。

自监督学习(SSL)推动了医学图像分析的发展,使模型能够从大量无标签数据中学习。然而,在脑部磁共振成像(MRI)中,多数3D模型仍局限于分割或分类某一特定任务,难以跨数据集、成像协议和下游任务泛化,限制了其临床应用价值。本文提出Volta-3D,一种基于3D视觉变压器的自监督框架,旨在学习可迁移的体积表示。该方法在学生-教师范式下联合对齐全局类别风格令牌与局部补丁令牌,并强制实现细粒度结构重建。这种全局-局部对齐策略有效应对脑部MRI语义多样性不足与解剖细节细微的问题,克服了现有SSL方法的局限。我们在多个分布外下游任务上评估了Volta-3D,包括海马体分割及性别、阿尔茨海默病与健康对照的分类。结果表明,其学习到的表征在所有任务上均优于随机初始化基线,展现出更强的迁移性与域偏移鲁棒性。因此,预训练阶段同时强化全局语义一致性和局部结构学习,有助于从无标签脑部MRI数据中实现更广泛的认知学习。整体而言,Volta-3D支持任务特定微调下的多任务高效表现,迈向通用且具备临床可行性的3D模型。

原文摘要 · Abstract (English)

Self-supervised learning (SSL) has advanced medical image analysis be enabling learning form large unlabelled data. However, in brain magnetic resonance imaging (MRI), most 3D models remain specialized for either segmentation of classification, limiting their ability to generalize across datasets, imaging protocols,, and downstream tasks. This lack of transferability constrains the clinical utility of 3D MRI models, despite the availability of unlabeled volumetric data. We present Volta-3D, a self-supervised 3D Vision Transformer framework designed to learn transferable volumetric representations. Volta-3D jointly aligns global class-style tokens and local patch tokens within a student-teacher paradigm and enforces fine-grained structural reconstruction. This combined global-local alignment addresses the limited semantic diversity and subtle anatomical characteristics of brain MRI, which challenges existing SSL approaches. We evaluate Volta-3D on multiple out-of-distribution downstream tasks, including hippocampal segmentation and classification of sex and Alzheimer's disease versus healthy controls. Across all tasks, representations learned by Volta-3D outperform randomly initialized baselines, demonstrating improved transferability and robustness under domain shift. Hence jointly enforcing global semantic consistency and local structural learning during pretraining enables broader concept learning from unlabeled brain MRI data. Overall VolTA-3D supports effective multi-task downstream performance with task-specific pertaining, a step towards generalizable and clinically viable 3D models.

自监督学习脑部MRI3D视觉变压器多任务迁移

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。