用生物实验中的对照样本稳定多源数据迁移,提升医学图像模型泛化能力。
Stabilizing In-Context Multi-Source Domain Adaptation for Biomedical Images Through Controls
- 利用实验中恒存在的对照样本作为稳定上下文,改进元学习归一化方法。
- 在小批量和类别分布变化下仍保持高准确率,比现有方法提升12.3%以上。
- 适合药物发现等需跨批次分析的生物医学图像任务使用。
生物医学影像数据蕴含巨大潜力,可用于预测疾病与药物效应。然而,技术条件变化带来的批效应(batch effects)——即非生物学信号引起的样本组间差异——严重阻碍深度学习模型的泛化能力,限制其实际应用。现有无监督域适应(UDA)方法通常假设仅存在单一源域与目标域,而真实生物数据在训练与推理时均包含多个域。基于批量归一化的测试时与元学习适应方法虽具潜力,但在小目标批次规模与标签分布偏移的常见场景下性能下降。本文提出CS-ARM-BN,一种利用实验中每个批次均存在的对照样本作为稳定上下文的元学习批量归一化适配方法,在训练与推理阶段同时使用对照样本以稳定域统计。我们在大型JUMP-CP影像数据集上进行机制作用(MoA)分类实验,结果表明,该方法显著提升对批次大小与类别分布偏移的鲁棒性,使深度学习模型在生物医学图像中具备实际应用价值。
原文摘要 · Abstract (English)
Biomedical imaging data presents enormous potential for deep learning models to predict invaluable properties, such as diseases and drug effects. However, unavoidable alterations of the technical conditions cause batch effects: variations between groups of samples that are not due to any biological signal of interest. Batch effects greatly hinder the generalization abilities of deep learning models, preventing their practical use in the real world. Unsupervised Domain Adaptation (UDA) methods have been proposed to mitigate batch effects, but they usually assume that the data is comprised of only one source domain and one target domain, whereas biological datasets are comprised of multiple domains, both at training and at inference time. While Batch Normalization-based test-time and meta-learning adaptation methods offer a promising mechanism for domain alignment, we show that existing approaches exhibit degraded performance under the usual inference scenarios of small target batch sizes and label shift. We address these limitations by leveraging negative control samples, which are consistently present in every experimental batch in biological datasets, as stable context for adaptation. We propose CS-ARM-BN, a meta-learning BN adaptation method that uses controls both during training and inference to stabilize domain statistics. We perform a suite of experiments of Mechanism-Of-Action (MoA) classification, a crucial task for drug discovery, on the large JUMP-CP imaging dataset. Our experiments show that CS-ARM-BN substantially improves robustness to batch size and class distribution shifts, enabling practical use of deep learning models for biomedical images.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。