arXiv:2607.22712cs.CVcs.AI2026-07

scMIR用图文对齐实现单细胞显微图像的统一表征,无需微调即可跨任务通用。

scMIR: a vision-language foundation model for single-cell light microscopy image representation

  • 结合自监督重建与文本引导对齐,统一编码形态与生物语义信息
  • 在16个基准数据集上超越现有通用与专用模型,泛化能力强
  • 适用于多种细胞类型、显微模态和实验条件下的高通量表型分析

单细胞光显微镜图像已成为表征细胞表型的重要数据源,但其复杂性和异质性给高通量自动化分析带来挑战。现有表征学习方法多依赖特定任务建模,受限于特定数据集和预设任务,难以跨不同细胞类型、显微模态和实验条件泛化。尽管近年通用方法提升了图像表征的泛化能力,但仍未能充分利用实验背景与生物上下文信息,制约复杂表型分析。本文提出scMIR,一种用于单细胞光显微镜图像表征的视觉-语言基础模型。通过协同结合自监督图像重建与文本引导的跨模态对齐,scMIR可在统一表示空间中同时编码形态学与生物语义信息。模型在207,957对图像-文本数据上预训练,涵盖多种细胞类型、显微模态及扰动条件。系统评估显示,scMIR在16个基准数据集上的各类复杂任务(包括细胞分类、聚类、表型推断与批次效应校正)中优于现有通用与任务导向方法。此外,scMIR展现出强跨任务泛化能力,无需任务特定微调。凭借独特优势,scMIR有望推动高通量表型分析流程的标准化与自动化,支持多样下游分析任务。

原文摘要 · Abstract (English)

Single-cell light microscopy images have become an important data source for characterizing cell phenotypes, but their complexity and heterogeneity pose challenges to high-throughput automated analysis. Existing representation learning methods mostly rely on task-oriented modeling, which is limited by specific datasets and predefined tasks, making them difficult to generalize across different cell types and microscopy modalities, and experimental conditions. Although general-purpose methods have improved the generalization ability of image representation in recent years, their limited utilization of experimental background and biological context information still poses challenges in complex phenotypic analysis. Here, we propose scMIR, a vision-language foundation model for single-cell light microscopy image representation. By synergistically combining self-supervised image reconstruction with text-guided cross-modal alignment, scMIR can simultaneously encode morphological and biological semantic information in a unified representation space. scMIR is pre-trained on 207,957 image-text pairs, covering various cell types, microscopy modalities, and perturbation conditions. scMIR outperforms existing general models and task-oriented methods as systematically evaluated on various complex tasks using 16 benchmark datasets, including cell classification, clustering, phenotype inference, and batch effect correction tasks. Furthermore, scMIR shows a strong generalization ability across various tasks without requiring task-specific fine-tuning. With its unique advantages, we envision scMIR may promote the standardization and automation of high-throughput phenotyping workflows through supporting various downstream analysis tasks.

单细胞视觉语言表型分析基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。