JUMP-lite让细胞图像表征模型可复现、低成本地对比测试。
JUMP-lite: Compact, reproducible benchmarking of cell representations

- 从115TB的JUMP数据中精选92GB子集,保留基因与化合物注释多样性。
- 压缩后仍能保持主要表型信号,五种模型性能差异明显可测。
- 配套开源框架Nahual确保实验可复现,适合科研人员快速评估模型。
基于图像的表型分析为药物发现和功能基因组学提供了丰富的表型特征。大型公共数据集如JUMP Cell Painting如今包含数百万张图像,可供系统研究。然而,仅JUMP数据就占115 TB,且评估方式碎片化,使多数研究者难以系统比较表征方法。为此,我们提出Nahual——一个用于可复现模型部署的开源框架,以及JUMP-lite,一个92.0 GB的子集,约为原始数据的1,250倍缩小,通过精心选择具有高置信度注释的扰动并采用有损JPEG XL压缩实现。利用这些资源,我们对五种表征方法进行了基准测试,包括传统特征(CellProfiler)和深度学习模型(MorphEM、OpenPhenom、SubCell、DINOv2)。适度压缩在保留信号方面表现良好。标准化的表型活性与一致性度量揭示了不同方法间的显著性能差异。JUMP-lite与Nahual共同构建了一个可访问、可复现的图像基细胞表征基准测试基础。
原文摘要 · Abstract (English)
Image-based profiling captures rich phenotypic signatures for drug discovery and functional genomics. Large public datasets like JUMP Cell Painting now provide millions of images for systematic study. However, JUMP alone occupies 115 TB, and fragmented evaluation practices make systematic comparisons of representation methods impractical for many researchers. Here we present Nahual, an open-source framework for reproducible model deployment, and JUMP-lite, a 92.0 GB subset of JUMP that is approximately 1,250-fold smaller, selected to cover genetic modalities and compound annotations and reduced via lossy JPEG XL compression. Using these resources, we benchmark five representation methods, including classical features (CellProfiler) and deep learning models (MorphEM, OpenPhenom, SubCell, DINOv2). Moderate compression broadly retains signal relative to uncompressed images. Standardized phenotypic activity and consistency metrics reveal meaningful performance differences across methods. Together, JUMP-lite and Nahual provide a foundation for accessible, reproducible benchmarking of image-based cell representations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。