新框架DITL提升乳腺钼靶图像分类,兼顾小样本与大规模数据。
Re-thinking Mammography Transfer Learning: The Dataset-Informed Transfer Learning (DITL) Framework for Breast Cancer Screening and Lesion Diagnosis

- 基于数据特征自适应加权和可学习间距的联合优化策略。
- 在大尺度数据集上准确率、F1和AUC均显著提升(p<0.0001)。
- 无需调参,计算开销低,适合临床筛查到诊断全流程应用。
提升钼靶图像分类性能在小规模精选数据集和大规模临床队列中仍是持续挑战。传统迁移学习方法常忽略数据集特异性,而近期邻域感知方法受限于窄任务且公式僵化,难以扩展至人群级数据。为此,我们提出数据驱动的迁移学习(DITL)框架,将数据导出的难度信号与基于邻域的三元组监督统一整合。DITL引入两个自适应组件:(i) 自适应难度加权交叉熵(A-DWCE),根据自监督特征空间中k近邻标签纯度为每样本分配权重;(ii) 自适应邻域表示三元组(A-NR-Triplet),使用可学习边距实现类内紧凑与类间分离。相比焦点损失,DITL无需超参数调优,去除启发式加权与固定边距,计算开销可忽略,形成鲁棒可扩展的优化策略。在大规模VinDR-Mammo数据集上,DITL在全图乳腺密度分类中达到最优性能,准确率、F1-score和AUC均有显著提升(p < 0.0001)。在小规模区域感兴趣(ROI)数据集上亦持续获得统计显著增益(p < 0.0001)。DITL贯通小样本病灶分析与大规模密度评估,建立了一种临床相关、可扩展且通用的钼靶分类框架,覆盖乳腺癌筛查到诊断全链条。
原文摘要 · Abstract (English)
Enhancing classification performance in mammography remains a persistent challenge across both small curated datasets and large-scale clinical cohorts. Conventional transfer learning approaches often neglect dataset-specific characteristics, while recent neighborhood-informed methods have been restricted to narrow tasks with rigid formulations, limiting their scalability to population-level datasets. To address these challenges, we propose the Dataset-Informed Transfer Learning (DITL) framework, which integrates dataset-derived difficulty signals with neighborhood-based triplet supervision in a unified objective. DITL introduces two adaptive components: (i) Adaptive Difficulty-Weighted Cross-Entropy (A-DWCE), which assigns per-sample weights based on k-nearest neighbor label purity in a self-supervised feature space, and (ii) Adaptive Neighborhood Representation Triplet (A-NR-Triplet), which enforces intra-class compactness and inter-class separation using a learnable margin. Unlike focal loss, DITL requires no hyperparameter tuning, removes heuristic weighting and fixed margins, and incurs negligible computational overhead, yielding a robust and scalable optimization strategy. On the large-scale VinDR-Mammo dataset, DITL achieves state-of-the-art performance for whole-image breast density classification, with significant improvements across accuracy, F1-score, and AUC (p < 0.0001). Beyond large cohorts, DITL also delivers consistent, statistically significant gains on small ROI datasets (p < 0.0001). By bridging small-scale lesion analysis with large-scale density estimation, DITL establishes a clinically relevant, scalable, and generalizable framework for mammography classification, spanning the full breast cancer screening-to-diagnosis spectrum.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。