用2D大模型语义提升3D脑动脉瘤分割精度,参数少且泛化强。
DINO-3DRA: Leveraging 2D Foundation Model Semantics for 3D Cerebral Aneurysm Segmentation

- 将冻结的DINOv3特征通过空间混合注入3D U-Net,保持解剖连续性
- 在多中心数据上达到Dice 0.758,比nnU-Net提升13%
- 无需微调即可避免灾难性失败,适合跨设备部署
3D旋转血管造影(3DRA)中动脉瘤精准分割受限于极端类别不平衡、与血管形态相似以及缺乏大规模3D预训练。2D视觉基础模型从17亿张图像中编码密集结构先验,但简单的逐切片迁移会破坏解剖连续性并导致优化不稳定。我们提出DINO-3DRA,一种双路径框架,通过Room-Lite空间混合与校准残差融合,将冻结的DINOv3特征有效注入3D U-Net主干网络。在多中心3DRA数据上,DINO-3DRA实现当前最优分割效果(Dice: 0.758;HD95: 2.75 mm),较nnU-Net提升13%,仅需572万可训练参数。消融实验表明,性能提升源于结构化的跨维度语义迁移,而非损失函数设计本身;注入的基础特征显著增强了动脉瘤与父血管间的解剖连贯性。在未微调CADA和SHINY-ICARUS数据集的情况下,DINO-3DRA消除所有基线架构中的灾难性失败案例,证明其在异构成像协议下具备强泛化能力。
原文摘要 · Abstract (English)
Accurate aneurysm segmentation in 3D rotational angiography (3DRA) is hindered by extreme class imbalance, morphological similarity to vessels, and absent large-scale 3D pretraining. 2D vision foundation models encode dense structural priors from 1.7 billion images, yet naïve slice-wise transfer fragments anatomical continuity and destabilises optimisation. We propose DINO-3DRA, a dual-path framework achieving effective cross-dimensional semantic transfer by injecting frozen DINOv3 features into a 3D U-Net backbone via Room-Lite spatial mixing and calibrated residual fusion. On multi-centre 3DRA data, DINO-3DRA achieves state-of-the-art aneurysm segmentation (Dice: 0.758; HD95: 2.75 mm; +13% over nnU-Net) with only 5.72M trainable parameters. Ablation studies confirm that gains arise from structured cross-dimensional transfer rather than loss design alone, with bridged foundation features improving anatomical continuity between aneurysms and parent vessels. Without fine-tuning on CADA and SHINY-ICARUS, DINO-3DRA eliminates all catastrophic failure cases observed in baseline architectures, demonstrating robust generalisation across heterogeneous imaging protocols.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。