构建首个面向大模型的虚拟细胞表型筛选基准,助力药物研发智能化。
AssayBench: An Assay-Level Virtual Cell Benchmark for LLMs and Agents

- 基于1920个公开CRISPR筛选数据,构建跨表型类别的表型预测任务
- 提出调整后nDCG指标,实现异质实验间性能连续对比
- 发现通用大模型在零样本下优于领域专用模型,可进一步优化提升
机器学习与大规模生物数据的进步推动了虚拟细胞建模的发展,有望加速生物学发现。其中最具前景的应用是体外表型筛选:模型预测细胞扰动在未见生物学情境下的效应。该任务融合异构文本输入与多样表型输出,特别适合大语言模型(LLMs)与智能体系统。然而,当前尚无标准基准,现有工作多聚焦于分子读数,与实际药物研发中的表型终点关联较弱。本文提出AssayBench,一个基于1920个公开CRISPR筛选数据、涵盖五类主要细胞表型的表型筛选预测基准。我们将筛选预测任务建模为每个实验的基因排序预测,并引入调整后的nDCG作为连续评估指标,以支持跨异质实验的性能比较。大量评估表明,现有方法仍远低于估算的性能上限,且零样本通用大模型表现优于生物学专用大模型及可训练基线。微调、集成与提示优化等技术可进一步提升大模型性能。总体而言,AssayBench为体外表型筛选与虚拟细胞模型的研究提供了一个实用的测试平台。
原文摘要 · Abstract (English)
Recent advances in machine learning and large-scale biological data collections have revived the prospect of building a virtual cell, a computational model of cellular behavior that could accelerate biological discovery. One of the most compelling promises of this vision is the ability to perform in silico phenotypic screens, in which a model predicts the effects of cellular perturbations in unseen biological contexts. This task combines heterogeneous textual inputs with diverse phenotypic outputs, making it particularly well-suited to LLMs and agentic systems. Yet, no standard benchmark currently exists for this task, as existing efforts focus on narrower molecular readouts that are only indirectly aligned with the phenotypic endpoints driving many real-world drug discovery workflows. In this work, we present AssayBench, a benchmark for phenotypic screen prediction, built from 1,920 publicly available CRISPR screens spanning five broad classes of cellular phenotypes. We formulate the screen prediction task as a gene rank prediction for each screen and introduce the adjusted nDCG, a continuous metric for comparing performance across heterogeneous assays. Our extensive evaluation shows that existing methods remain far from empirically estimated performance ceilings and zero-shot generalist LLMs outperform biology-specific LLMs and trainable baselines. Optimization techniques such as fine-tuning, ensembling, and prompt optimization can further improve LLM performance on this task. Overall, AssayBench offers a practical testbed for measuring progress toward in silico phenotypic screening and, more broadly, virtual cell models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。