arXiv:2608.24364cs.CV2026-08

通过削弱全局语义对齐,提升医学影像对细粒径解剖结构的感知能力。

B-MIM: Biased Masked Image Modeling for Generalizable Segmentation of Fine-Grained Anatomical Structures

论文配图:B-MIM: Biased Masked Image Modeling for Generalizable Segmentation of Fine-Grained Anatomical Structures
图 1 · 摘自论文原文
  • 设计B-MIM,随机降低全局对齐,专注局部图像重建。
  • 在9,955例腹部CT上预训练3D Swin Transformer,提升血管拓扑精度。
  • 仅微调少量参数,即可实现优于全微调的肿瘤分割效果。

自监督预训练能生成可迁移的医学影像表征,但多数CT编码器仍偏向粗粒度语义理解,难以捕捉血管或小肿瘤等细粒径解剖结构。本文提出有偏掩码图像建模(B-MIM),对iBOT目标进行改进,通过随机降低全局语义对齐,优先关注局部块重建。该偏差促使编码器捕捉高频形态细节与结构连续性。我们构建了来自17个公开来源的9,955例多机构腹部CT数据集,并用B-MIM预训练3D Swin Transformer骨干网络。在跨数据集的肝脏血管分割实验中,所提编码器在拓扑保真度(clDice)上表现更优,且在肿瘤分割任务中达到与全微调基线相当的Dice分数,仅需更新部分参数。结果表明,预训练阶段降低全局语义压力有助于提升对复杂解剖结构的泛化能力。

原文摘要 · Abstract (English)

Self-supervised pretraining enables transferable representations for medical imaging, yet most CT encoders remain biased toward coarse semantic understanding, limiting their sensitivity to fine-grained anatomical structures such as vessels or small tumors. In this paper, we introduce Biased Masked Image Modeling (B-MIM), a modification of the iBOT objective that stochastically reduces global semantic alignment to prioritize local patch reconstruction. This bias encourages the encoder to capture high-frequency morphological details and structural continuity. We curate a multi-institutional CT abdominal dataset of 9,955 filtered studies from 17 public sources and pretrain a 3D Swin Transformer backbone using B-MIM. Across inter-dataset experiments on liver vessel segmentation, the proposed encoder improves topological fidelity (clDice) and achieves competitive Dice scores in tumor segmentation, compared to fully fine-tuned baselines, despite updating only a fraction of the parameters. Our results suggest that reducing global semantic pressure during pretraining enhances generalization to intricate anatomical structures.

医学影像自监督学习细粒径分割3D Swin

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。