用对抗生成方法提升细胞筛选数据的批次鲁棒性
Adversarial Batch Representation Augmentation for Batch Correction in High-Content Cellular Screening
- 将批次校正视为域泛化问题,通过参数化特征统计建模批次波动
- 在表示空间合成最坏情况批次扰动,保持细粒度分类能力
- 适用于无先验信息的未知批次场景,尤其适合高通量细胞筛选
高通量细胞筛选常生成大量细胞着色图像用于表型分析,但不同实验批次间的技术差异会引入生物批次效应,导致协变量偏移,降低深度学习模型在未见数据上的泛化能力。现有批次校正方法通常依赖额外先验信息(如处理或培养条件)或难以推广到未知生物批次。本文将生物批次缓解视为域泛化问题,提出对抗批次表示增强(ABRA)。ABRA通过参数化特征统计来显式建模批次间统计波动,并在最小-最大优化框架下,基于严格的角几何边界,在表示空间主动合成最坏情况的生物批次扰动,以保持细粒度类别可区分性。为防止对抗探索中的表示坍缩,引入协同分布对齐目标。在大规模RxRx1和RxRx1-WILDS基准上的广泛评估表明,ABRA在siRNA扰动分类任务上达到新最优性能。
原文摘要 · Abstract (English)
High-Content Screening routinely generates massive volumes of cell painting images for phenotypic profiling. However, technical variations across experimental executions inevitably induce biological batch (bio-batch) effects. These cause covariate shifts and degrade the generalization of deep learning models on unseen data. Existing batch correction methods typically rely on additional prior knowledge (e.g., treatment or cell culture information) or struggle to generalize to unseen bio-batches. In this work, we frame bio-batch mitigation as a Domain Generalization (DG) problem and propose Adversarial Batch Representation Augmentation (ABRA). ABRA explicitly models batch-wise statistical fluctuations by parameterizing feature statistics as structured uncertainties. Through a min-max optimization framework, it actively synthesizes worst-case bio-batch perturbations in the representation space, guided by a strict angular geometric margin to preserve fine-grained class discriminability. To prevent representation collapse during this adversarial exploration, we introduce a synergistic distribution alignment objective. Extensive evaluations on the large-scale RxRx1 and RxRx1-WILDS benchmarks demonstrate that ABRA establishes a new state-of-the-art for siRNA perturbation classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。