arXiv:2503.00744cs.CVcs.AI2025-03被引 2

针对医学影像微调中的混杂变量问题,提出高效数据筛选方法。

Confounder-Aware Medical Data Selection for Fine-Tuning Pretrained Vision Models

  • 识别数据中混杂变量,基于距离选择代表性样本
  • 在有限数据量下提升模型微调效果,减少偏差影响
  • 适合关注数据质量与模型泛化能力的研究者

大规模预训练视觉基础模型的出现推动了医学影像领域的发展。然而,在下游微调中选择合适数据仍面临标注成本高、隐私问题及混杂变量干扰等挑战。本文提出一种考虑混杂变量的医学数据筛选方法,旨在通过策略性缓解混杂变量的负面影响,同时保持数据集自然分布,以最小数据量选出最具代表性的样本。方法首先识别数据中的混杂变量,再设计基于距离的数据选择策略,在数据规模受限条件下实现有约束的采样。在多种医学影像模态上的大量实验验证了该方法的有效性,相较于其他数据选择方法,显著降低了混杂变量的影响,提升了微调效率。

原文摘要 · Abstract (English)

The emergence of large-scale pre-trained vision foundation models has greatly advanced the medical imaging field through the pre-training and fine-tuning paradigm. However, selecting appropriate medical data for downstream fine-tuning remains a significant challenge considering its annotation cost, privacy concerns, and the detrimental effects of confounding variables. In this work, we present a confounder-aware medical data selection approach for medical dataset curation aiming to select minimal representative data by strategically mitigating the undesirable impact of confounding variables while preserving the natural distribution of the dataset. Our approach first identifies confounding variables within data and then develops a distance-based data selection strategy for confounder-aware sampling with a constrained budget in the data size. We validate the superiority of our approach through extensive experiments across diverse medical imaging modalities, highlighting its effectiveness in addressing the substantial impact of confounding variables and enhancing the fine-tuning efficiency in the medical imaging domain, compared to other data selection approaches.

医学影像数据筛选混杂变量微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。