arXiv:2411.02572cs.LGcs.AI2024-11ICML被引 18

训练19亿参数模型,让细胞显微图像表征更一致,提升药物研发效率。

ViTally Consistent: Scaling Biological Representation Learning for Cell Microscopy

  • 用80亿图像块预训练1.9亿参数ViT-G/8模型,提升生物表征一致性。
  • 基因扰动线性可分性提升60%,全基因组关系召回率创纪录。
  • 通过生物导向的线性探测,发现中间层比最终层更优,适合生物分析。

大规模细胞显微镜筛选用于药物发现和分子生物学研究,以分析数百万种化学与基因扰动对细胞的影响。为支持下游分析,需建立能将图像映射到稳定表征空间的模型,使具有相似生物学效应的扰动产生相似表示。本文提出迄今最大的细胞显微图像基础模型:基于80亿图像块训练的19亿参数ViT-G/8 MAE。相比先前发布的ViT-L/8 MAE,该模型在基因扰动线性可分性上提升60%,并在全基因组生物关系召回与重复实验一致性基准上表现最佳。此外,我们开发了两项关键方法:(1) 使用经过筛选且多样化的数据集进行训练;(2) 采用生物启发的线性探测任务,在每个Transformer块中搜索最适合全基因组筛选的表示。研究发现,许多自监督视觉变换器(无论在自然图像或显微图像上预训练)在中间层的表征比常用最终层更具生物学意义。本研究为构建大规模生物数据基础模型提供了通用策略与洞见。

原文摘要 · Abstract (English)

Large-scale cell microscopy screens are used in drug discovery and molecular biology research to study the effects of millions of chemical and genetic perturbations on cells. To use these images in downstream analysis, we need models that can map each image into a feature space that represents diverse biological phenotypes consistently, in the sense that perturbations with similar biological effects have similar representations. In this work, we present the largest foundation model for cell microscopy data to date, a new 1.9 billion-parameter ViT-G/8 MAE trained on over 8 billion microscopy image crops. Compared to a previous published ViT-L/8 MAE, our new model achieves a 60% improvement in linear separability of genetic perturbations and obtains the best overall performance on whole-genome biological relationship recall and replicate consistency benchmarks. Beyond scaling, we developed two key methods that improve performance: (1) training on a curated and diverse dataset; and, (2) using biologically motivated linear probing tasks to search across each transformer block for the best candidate representation of whole-genome screens. We find that many self-supervised vision transformers, pretrained on either natural or microscopy images, yield significantly more biologically meaningful representations of microscopy images in their intermediate blocks than in their typically used final blocks. More broadly, our approach and results provide insights toward a general strategy for successfully building foundation models for large-scale biological data.

生物表征视觉变换器自监督学习细胞成像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。