DINOv3在512像素分辨率下对成人胸片分类效果最佳,尤其搭配ConvNeXt-B模型。
Resolution scaling governs DINOv3 transfer performance in chest radiograph classification
- 通过高分辨率适配与梯度锚定自蒸馏提升视觉表征能力
- 512×512分辨率下,对小病灶和边界依赖性异常提升显著
- 适合关注医学影像迁移学习性能的临床研究者
自监督学习(SSL)虽提升了视觉表征能力,但在胸部放射影像中的价值仍不明确。DINOv3通过梯度锚定自蒸馏和显式高分辨率适配扩展了早期模型。我们基于7个包含816,183张胸片的儿童与成人数据集,对比了DINOv3、DINOv2及监督ImageNet初始化表现。评估了ViT-B/16与ConvNeXt-B在224像素和512像素全微调下的性能,并在三个队列中进行了1024像素实验。还分析了参数高效适应、合成标签污染、外部验证、冻结7B特征及计算效率。主要指标为多标签平均AUROC。在成人队列中,DINOv3在224×224时未稳定优于DINOv2,但在512×512时成为最强初始化,尤其配合ConvNeXt-B;小病灶和边界依赖性异常改善最明显,大结构异常变化不大。儿童队列未见DINOv3、更高分辨率或主干网络选择的显著优势。1024×1024极少提升性能且大幅增加计算开销。无论全微调或参数高效适应,ConvNeXt-B始终优于ViT-B/16。外部验证保持512×512下DINOv3的优势,而合成标签污染表明该优势非单纯噪声鲁棒性所致。对于成人胸片分类,DINOv3在512×512分辨率下表现最优,尤其是搭配ConvNeXt-B。512×512全微调中等规模模型在性能与成本间取得最佳平衡。
原文摘要 · Abstract (English)
Self-supervised learning (SSL) has improved visual representation learning, but its value in chest radiography remains uncertain. DINOv3 extends earlier SSL models through Gram-anchored self-distillation and explicit high-resolution adaptation. Whether these changes improve transfer learning for chest radiograph classification has not been established. We benchmarked DINOv3 against DINOv2 and supervised ImageNet initialization across seven chest radiograph datasets comprising 816,183 radiographs from pediatric and adult cohorts. ViT-B/16 and ConvNeXt-B were evaluated under full fine-tuning at 224 and 512 pixels, with targeted 1024 experiments on three cohorts. Additional analyses examined parameter-efficient adaptation, synthetic label corruption, external validation, frozen 7B features, and computational efficiency. The primary outcome was mean AUROC across labels. In adult cohorts, DINOv3 did not consistently outperform DINOv2 at 224 x 224 pixels, but became the strongest initialization at 512 x 512, especially with ConvNeXt-B. Gains were greatest for small focal and boundary-dependent abnormalities, whereas large-structure findings changed little. The pediatric cohort showed no significant benefit from DINOv3, higher resolution, or backbone choice. Scaling to 1024 x 1024 rarely improved performance and markedly increased computational cost. ConvNeXt-B remained superior to ViT-B/16 under both full and parameter-efficient adaptation. External validation preserved the 512 x 512 DINOv3 advantage, whereas synthetic label corruption showed that this benefit should not be interpreted simply as superior noise robustness. For adult chest radiograph classification, DINOv3 provides its most reliable benefit at 512 x 512 pixels, particularly with ConvNeXt-B. Fully adapted mid-sized models at 512 x 512 pixels provided the best performance-cost trade-off in our benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。