利用目标域类别比例约束伪标签,提升医学图像域适应性能
Weakly-Supervised Domain Adaptation with Proportion-Constrained Pseudo-Labeling
- 基于目标域类别比例生成约束伪标签,无需额外标注
- 在仅5%标签情况下优于半监督域自适应方法
- 对噪声比例信息仍具鲁棒性,适合真实医疗场景
领域偏移是机器学习中的重大挑战,尤其在医疗应用中,不同机构因数据采集方式、设备和流程差异导致数据分布不同,使源域训练的模型在目标域上性能下降。现有域适应方法大多难以应对源域与目标域类别比例不同的情况。本文提出一种弱监督域适应方法,利用目标域类别比例信息(通常可通过先验知识或统计报告获得),基于比例约束为无标签目标数据分配伪标签,从而提升性能而无需额外标注。在两个内窥镜数据集上的实验表明,该方法在仅5%目标域样本有标签时仍优于半监督域适应技术。此外,使用噪声比例标签的实验也验证了方法的鲁棒性,进一步证明其在实际应用场景中的有效性。
原文摘要 · Abstract (English)
Domain shift is a significant challenge in machine learning, particularly in medical applications where data distributions differ across institutions due to variations in data collection practices, equipment, and procedures. This can degrade performance when models trained on source domain data are applied to the target domain. Domain adaptation methods have been widely studied to address this issue, but most struggle when class proportions between the source and target domains differ. In this paper, we propose a weakly-supervised domain adaptation method that leverages class proportion information from the target domain, which is often accessible in medical datasets through prior knowledge or statistical reports. Our method assigns pseudo-labels to the unlabeled target data based on class proportion (called proportion-constrained pseudo-labeling), improving performance without the need for additional annotations. Experiments on two endoscopic datasets demonstrate that our method outperforms semi-supervised domain adaptation techniques, even when 5% of the target domain is labeled. Additionally, the experimental results with noisy proportion labels highlight the robustness of our method, further demonstrating its effectiveness in real-world application scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。