用大模型筛选相似药物,预测新药在细胞中的基因反应。
LLM-Guided Retrieval for Prediction of Molecular Perturbation Responses

- 用大模型从已测药物中筛选生物相关候选药
- 在未见细胞系上提升预测相关性与准确性
- 适合药物研发中零样本预测场景
预测小分子扰动在不同细胞系中的转录组响应是药物发现的核心挑战,但全面测试药物-细胞组合不可行。本文将分子扰动预测建模为检索与聚合:通过聚合少量生物学相关的已测化合物响应,近似未测量药物在目标细胞系中的效应。提出LLM-Guided Retrieval(LGR),利用大语言模型(LLM)对目标细胞系中已测药物进行排序,再以固定均值聚合其表达变化量生成预测。在Tahoe-100M单细胞扰动图谱上评估,涵盖未见药物、未见细胞系及开放世界三种情形。LGR在所有设置下均优于药物均值、ChemCPA和基于化学结构的kNN基线,尤其在未见细胞系泛化上表现突出,相关性更高且误差更低。同时提升基因调控方向(符号)准确率,表明即使在幅度指标相近时,仍更好恢复了生物学有意义的扰动效应。结果表明,检索质量而非预测器复杂度,是零样本分子扰动预测的关键驱动因素,且大模型作为受限检索模块可提供有效生物先验。
原文摘要 · Abstract (English)
Predicting transcriptomic responses to small-molecule perturbations across cell lines is central to drug discovery, but exhaustive profiling of drug-cell combinations is infeasible. We frame molecular perturbation prediction as retrieve-and-aggregate: approximate an unmeasured drug's response in a cell line by aggregating measured responses of a small set of biologically related compounds. We propose LLM-Guided Retrieval (LGR), where a large language model (LLM) ranks candidate neighbor drugs (restricted to those profiled in the target cell line); after which a fixed mean aggregator combines their observed expression deltas to form the prediction. We evaluate on the Tahoe-100M single-cell perturbation atlas under unseen-drug, unseen-cell-line, and open-world regimes. LGR consistently improves over drug mean, ChemCPA, and chemistry-based kNN baselines, with the strongest gains for unseen cell-line generalization, where it achieves higher correlation and lower error than mean baselines. Across settings, LGR improves directional (sign) accuracy of gene regulation, indicating better recovery of biologically meaningful perturbation effects even when magnitude-based metrics are similar. These results suggest that retrieval quality, rather than predictor complexity, is a key driver of zero-shot molecular perturbation prediction, and that LLMs can provide a useful biological prior when used as constrained retrieval modules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。