用知识图谱+强化学习,高效预测基因敲除对基因表达的影响。
Knowledge Graphs and Reasoning LLMs for Finding Simple Yet Effective Transcriptomic Perturbation Predictors

- 基于知识图谱的K近邻法直接预测基因扰动效果。
- 该方法在未见扰动上表现优于多数现有模型。
- 强化学习微调的LLM可提升泛化能力,适合生物医学预测场景。
预测未知基因敲除对转录组表达的影响仍是虚拟细胞模型的重大挑战。近期研究通过生物知识图谱提供相似扰动的概念,提升了训练扰动外的外推能力。本文表明,最简单的基于知识图谱的K近邻方法在此任务上已具高度竞争力,进一步通过强化学习(RL)优化推理型大模型调整邻域结构,性能可达到当前最优水平。实验显示,该方法在Replogle等(2022)细胞系上的预测表现与先进方法相当,且即使未直接训练于差异表达任务,强化学习仍显著提升其下游表现。结果证明知识图谱作为模型先验的有效性,并揭示强化学习可将大模型转化为可泛化的复杂生物响应预测工具。
原文摘要 · Abstract (English)
Predicting the effect of an unseen gene knockout perturbation on transcriptomic gene expression remains a highly challenging problem for virtual cell models. Recent progress has been made by leveraging biological knowledge graphs to provide a notion of similar perturbation, allowing for improved extrapolation beyond the set of training perturbations. In this work, we demonstrate that the simplest model to leverage these assumptions - a K-nearest neighbour from the knowledge graph - achieves highly competitive performance on this task, and that this can be improved further using LLMs optimised via reinforcement learning (RL) for predictive performance. Specifically, we find that the K-nearest neighbour approach beats almost all methods on out-of-distribution perturbation prediction, and when a reasoning LLM is trained via RL to make changes to the neighbourhood, it obtains equivalent performance to current state of the art methods on the cell lines from Replogle et al. (2022). We also demonstrate that the RL training improves the LLM's performance on the downstream task of differential expression prediction, despite not being trained on this directly. Overall, these findings demonstrate the efficacy of knowledge graphs as model priors, and show early signs that RL can refine LLMs into generalizable tools for predicting complex biological responses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。