arXiv:2509.19896cs.CV2025-09被引 3

用跨孔对齐提升细胞图像表征,更省数据、更抗批次效应。

Efficient Cell Painting Image Representation Learning via Cross-Well Aligned Masked Siamese Network

  • 通过跨孔对齐的掩码孪生网络,让相同扰动的细胞图像特征一致。
  • 在基因关系检索中性能超现有方法,数据量少10倍、参数少67倍。
  • 适合小样本、低算力场景下的细胞表型建模与药物研发加速。

预测细胞对化学和基因扰动的表型响应的计算模型可加速药物发现,减少昂贵的湿实验迭代。然而,提取生物有意义且批次鲁棒的细胞着色表征仍具挑战。传统自监督和对比学习方法通常需要大规模模型和大量精心标注的数据,仍受批次效应影响。本文提出跨孔对齐掩码孪生网络(CWA-MSN),一种新型表征学习框架,通过将同一扰动下不同孔中细胞的嵌入对齐,确保语义一致性,缓解批次效应。该方法融入掩码孪生架构,在保持数据和参数高效的同时,捕捉细微形态特征。例如,在基因-基因关系检索基准测试中,CWA-MSN 相较于公开的自监督方法 OpenPhenom 和对比学习方法 CellCLIP,分别提升 +29% 和 +9% 的得分,且训练仅需 0.2M 张图像(相较 OpenPhenom 的 2.2M)或 22M 参数(相较 CellCLIP 的 1.48B)。大量实验表明,CWA-MSN 是一种简单有效的细胞图像表征学习方法,即使在数据和参数受限条件下也能实现高效表型建模。

原文摘要 · Abstract (English)

Computational models that predict cellular phenotypic responses to chemical and genetic perturbations can accelerate drug discovery by prioritizing therapeutic hypotheses and reducing costly wet-lab iteration. However, extracting biologically meaningful and batch-robust cell painting representations remains challenging. Conventional self-supervised and contrastive learning approaches often require a large-scale model and/or a huge amount of carefully curated data, still struggling with batch effects. We present Cross-Well Aligned Masked Siamese Network (CWA-MSN), a novel representation learning framework that aligns embeddings of cells subjected to the same perturbation across different wells, enforcing semantic consistency despite batch effects. Integrated into a masked siamese architecture, this alignment yields features that capture fine-grained morphology while remaining data- and parameter-efficient. For instance, in a gene-gene relationship retrieval benchmark, CWA-MSN outperforms the state-of-the-art publicly available self-supervised (OpenPhenom) and contrastive learning (CellCLIP) methods, improving the benchmark scores by +29\% and +9\%, respectively, while training on substantially fewer data (e.g., 0.2M images for CWA-MSN vs. 2.2M images for OpenPhenom) or smaller model size (e.g., 22M parameters for CWA-MSN vs. 1.48B parameters for CellCLIP). Extensive experiments demonstrate that CWA-MSN is a simple and effective way to learn cell image representation, enabling efficient phenotype modeling even under limited data and parameter budgets.

细胞表型自监督学习图像表征药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。