arXiv:2602.13930cs.CVcs.LG2026-02

用低分辨率乳腺影像实现与顶尖模型相当的3年癌变风险预测。

MamaDino: A Hybrid Vision Model for Breast Cancer 3-Year Risk Prediction

  • 融合卷积网络与视觉变压器,显式建模双侧乳腺不对称性。
  • 在512x512分辨率下达到AUC 0.736(内分布)和0.677(外分布)。
  • 适合关注可解释性与计算效率的医学影像风险预测研究者。

乳腺癌筛查正从统一间隔转向个性化风险适配策略。深度学习已使基于图像的风险模型在1至5年预测上超越传统临床模型,但主流系统(如Mirai)通常使用卷积主干、超高分辨率输入(>100万像素)和简单多视图融合,且对对侧不对称性建模有限。我们假设:结合卷积与变压器的互补归纳偏置,并显式建模对侧不对称性,可在更低分辨率下仍达到顶尖性能,表明以更结构化方式处理低细节图像亦可恢复高精度。本文提出MamaDino,一种面向乳腺摄影的多视图注意力式DINO模型。MamaDino在512x512分辨率下融合冻结的自监督DINOv3 ViT-S特征与可训练CNN编码器,并通过BilateralMixer聚合双侧乳腺信息,输出3年乳腺癌风险评分。在53,883名来自OPTIMAM(英国)女性数据上训练,并在匹配的3年病例对照队列中评估:四个筛查中心的内分布测试集及一个未见站点的外部外分布队列。在乳腺级水平,MamaDino在内部和外部测试中均匹配Mirai表现,而输入像素仅约其1/13。引入BilateralMixer后,内分布AUC提升至0.736(原0.713),外分布达0.677(原0.666),且在年龄、种族、扫描仪、肿瘤类型和分级上均保持稳定表现。结果表明,显式对侧建模与互补归纳偏置使模型在低分辨率下仍能实现与Mirai相当的预测性能。

原文摘要 · Abstract (English)

Breast cancer screening programmes increasingly seek to move from one-size-fits-all interval to risk-adapted and personalized strategies. Deep learning (DL) has enabled image-based risk models with stronger 1- to 5-year prediction than traditional clinical models, but leading systems (e.g., Mirai) typically use convolutional backbones, very high-resolution inputs (>1M pixels) and simple multi-view fusion, with limited explicit modelling of contralateral asymmetry. We hypothesised that combining complementary inductive biases (convolutional and transformer-based) with explicit contralateral asymmetry modelling would allow us to match state-of-the-art 3-year risk prediction performance even when operating on substantially lower-resolution mammograms, indicating that using less detailed images in a more structured way can recover state-of-the-art accuracy. We present MamaDino, a mammography-aware multi-view attentional DINO model. MamaDino fuses frozen self-supervised DINOv3 ViT-S features with a trainable CNN encoder at 512x512 resolution, and aggregates bilateral breast information via a BilateralMixer to output a 3-year breast cancer risk score. We train on 53,883 women from OPTIMAM (UK) and evaluate on matched 3-year case-control cohorts: an in-distribution test set from four screening sites and an external out-of-distribution cohort from an unseen site. At breast-level, MamaDino matches Mirai on both internal and external tests while using ~13x fewer input pixels. Adding the BilateralMixer improves discrimination to AUC 0.736 (vs 0.713) in-distribution and 0.677 (vs 0.666) out-of-distribution, with consistent performance across age, ethnicity, scanner, tumour type and grade. These findings demonstrate that explicit contralateral modelling and complementary inductive biases enable predictions that match Mirai, despite operating on substantially lower-resolution mammograms.

乳腺癌风险预测多模态融合低分辨率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。