模型预训练数据比训练目标更能决定隐私保护下的医疗影像诊断效果
The pretraining domain outweighs the training objective in setting the privacy-utility trade-off of differentially private medical image analysis
- 独立控制预训练目标与领域,对比不同初始化对隐私-效用权衡的影响
- 胸部影像监督预训练在24/25组合中表现最优,隐私预算越紧优势越大
- 私有化预训练数据仅损失约5分,仍优于公开数据初始化
差分隐私可保护患者图像,但会降低诊断准确率,而模型初始化是目前最有效的缓解方法。实践中越来越多使用大规模通用自监督编码器,但现有研究中预训练目标与领域混杂,无法区分哪个因素在隐私条件下更关键,且常假设预训练语料为公开数据,即使其中包含患者图像。本研究使用不同初始化(独立改变目标与领域)的ConvNeXt分类器,在四种隐私预算下及无隐私设置中,通过差分隐私随机梯度下降训练,并在来自四个国家的五个外部数据集、超过59万张胸部X光片上进行本地评估。结果显示,基于胸部影像的监督预训练在24/25种数据集与预算组合中排名第一。其相对于ImageNet的性能提升从2.5增至14.6点(宏平均AUC),领域影响比目标影响大2.2至3.4倍。私有化预训练语料导致约5分损失,但在隐私条件下仍优于所有公开初始化。低秩适配消除约一半残余差距,且域内预训练显著提升最差人群子组表现。在隐私约束下,模型预训练的数据领域比训练方式更重要。
原文摘要 · Abstract (English)
Differential privacy protects the patients whose images train medical imaging models, but it lowers diagnostic accuracy, and the initialization is the strongest known remedy. Practice increasingly favors large generic self-supervised encoders. Yet the pretraining objective and the pretraining domain are confounded in existing comparisons, so which one preserves utility under privacy is unknown, and the pretraining corpus is treated as public even when it holds patient images. We trained ConvNeXt classifiers with differentially private stochastic gradient descent from five initializations that vary the objective and the domain independently, at four privacy budgets and without privacy, and evaluated them locally on more than 590,000 chest radiographs from five external datasets in four countries. Supervised pretraining on chest radiographs ranked first in 24 of 25 dataset and budget combinations. Its lead over ImageNet grew from 2.5 to 14.6 points of macro-averaged area under the receiver operating characteristic curve as the budget tightened, and the domain effect exceeded the objective effect by a factor of 2.2 to 3.4. Pretraining that corpus privately cost about 5 points and, under privacy, still beat every public initialization. Low-rank adaptation removed about half the residual gap, and in-domain pretraining raised the worst-performing demographic subgroup. Under privacy, what a model was pretrained on outweighs how it was pretrained.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。