构建可跨数据集泛化的细胞形态表征模型,免微调即可应对新数据挑战。
CellPainTR: Generalizable Representation Learning for Cross-Dataset Cell Painting Analysis
- 基于Transformer架构,引入源特定上下文标记,学习抗批次效应的细胞形态表示。
- 在JUMP数据集上优于ComBat和Harmony,在未见数据集上仍保持高性能。
- 适合需要跨研究整合图像数据的生物医学研究人员使用。
大规模生物发现需要整合如JUMP Cell Painting联盟的海量异构数据,但技术性批次效应及缺乏可泛化模型仍是主要障碍。为此,我们提出CellPainTR,一种基于Transformer的架构,旨在学习对批次效应鲁棒的细胞形态基础表示。与传统方法需在新数据上重新训练不同,CellPainTR通过源特定上下文标记设计,实现无需微调即可有效泛化至完全未见数据集。我们在大规模JUMP数据集上验证了其性能,结果表明其在批次整合与生物信号保留方面均优于ComBat和Harmony。关键的是,在未见的Bray et al.数据集上进行具有挑战性的分布外(OOD)任务测试时,其依然保持高表现,克服了显著的领域与特征偏移。本工作为基于图像的分子表型分析奠定了基础模型,推动更可靠、可扩展的跨研究生物分析。
原文摘要 · Abstract (English)
Large-scale biological discovery requires integrating massive, heterogeneous datasets like those from the JUMP Cell Painting consortium, but technical batch effects and a lack of generalizable models remain critical roadblocks. To address this, we introduce CellPainTR, a Transformer-based architecture designed to learn foundational representations of cellular morphology that are robust to batch effects. Unlike traditional methods that require retraining on new data, CellPainTR's design, featuring source-specific context tokens, allows for effective out-of-distribution (OOD) generalization to entirely unseen datasets without fine-tuning. We validate CellPainTR on the large-scale JUMP dataset, where it outperforms established methods like ComBat and Harmony in both batch integration and biological signal preservation. Critically, we demonstrate its robustness through a challenging OOD task on the unseen Bray et al. dataset, where it maintains high performance despite significant domain and feature shifts. Our work represents a significant step towards creating truly foundational models for image-based profiling, enabling more reliable and scalable cross-study biological analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。