评估医学影像无监督域适应全流程,发现选模型难题制约临床落地。
How Far from Clinical Deployment? Evaluating the Complete Unsupervised Domain Adaptation Pipeline in Medical Imaging

- 系统测试11个跨域场景下10种算法与13种无标签选型方法
- 8万+模型评估显示选型方法普遍落后最优模型,差距显著且结构性
- 集成与少量标注可缩小差距,但无法完全解决选型困局
将无监督域适应(UDA)部署于临床实践需同时决定使用哪个算法及选择哪个训练好的模型。然而目标域无标签,无法直接评估模型性能,导致难以抉择。本文通过联合评估完整的UDA流程,涵盖域适应与无标签选型两步。研究覆盖来自9个医学影像数据集的11个临床相关跨域场景,对比10种UDA算法和13种无标签选型方法(验证器),共评估超过80,000个训练模型。结果表明:虽通常存在表现优异的适配模型,但无标签条件下识别它非常困难——验证器选出的模型在目标域上性能显著低于最优可用模型,且无任何验证器表现出持续可靠性。为缩小差距,我们探索了集成与少量目标标签预算两种策略,两者均有效减小性能差,但未彻底解决。总体而言,可部署的UDA依赖完整流程;解决尚未充分研究的选型环节,或能大幅推进当前UDA向临床应用靠近。
原文摘要 · Abstract (English)
Deploying unsupervised domain adaptation (UDA) in clinical practice requires choosing which algorithm to use and which of its trained models to ship. However, the deployment (target) domain is unlabeled, so models cannot be evaluated directly on it, leaving it unclear which to select. We address this by evaluating the complete UDA pipeline, considering both adaptation and label-free selection together. Our study covers eleven clinically relevant cross-domain scenarios from nine medical imaging datasets, with ten UDA algorithms and 13 label-free selection methods (validators), evaluating over 80,000 trained models in total. By this, we find that a capable adapted model usually exists, but identifying it without target labels is difficult: the validator-selected models leave a large and structural target performance gap to the best available one, with no evaluated validator consistently reliable. Towards closing it, we explore two strategies, ensembling and a small target-labeling budget; both narrow this gap but do not close it entirely. Overall, deployable UDA depends on the complete pipeline; addressing the less explored selection step could bring much of current UDA closer to clinical use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。