arXiv:2606.01042cs.LGcs.AI2026-06被引 2

用对比证据提升大模型对基因扰动的预测能力

Plausibility Is Not Prediction: Contrastive Evidence for LLM-Based Cellular Perturbation Reasoning

论文配图:Plausibility Is Not Prediction: Contrastive Evidence for LLM-Based Cellular Perturbation Reasoning
图 1 · 摘自论文原文
  • 将扰动预测重构为对比任务,通过正负例对照增强推理
  • 在药物扰动数据上使大模型性能提升28.6%,基因级准确率达0.703
  • 适合生物信息学与AI制药领域研究者参考

扰动实验是理解细胞机制的核心,但成本高且数据稀疏,促使人们尝试预测未观测条件下的基因表达响应。近期方法利用大语言模型(LLM)作为“虚拟细胞”模拟器,通过逐步、基于知识的机制推理推断差异表达,展现出可解释、知识驱动的潜力。然而我们发现:生物合理性不等于预测准确性——这些方法虽生成合理解释,却系统性高估差异表达,整体表现常低于简单基因频率基线,且单基因层面退化至随机水平,暴露出对基因固有响应倾向的依赖。问题根源在于证据呈现方式:现有方法孤立评估扰动-基因对,未揭示相关扰动在同一基因上的差异效应。为此,我们提出CORE(对比关系证据组织),将预测重构为比较任务,利用生物医学知识图谱检索关联扰动的正负例证据。在多种设置下,CORE显著提升校准度与特异性:例如,在药物扰动数据上,CORE-Reasoning使Qwen3.5-9B的聚合指标提升28.6%;在通用扰动数据上,CORE-Voting将平均四细胞系的宏观基因级AUROC从随机水平提升至0.703。这表明对比证据组织对可靠的大模型扰动推理至关重要。

原文摘要 · Abstract (English)

Perturbation experiments are central to understanding cellular mechanisms, but remain costly and sparse, motivating prediction of gene expression responses for unobserved conditions. A promising recent direction leverages large language models (LLMs) as "virtual cell" simulators-using stepwise, knowledge-grounded mechanistic reasoning to infer differential expression-pointing toward an interpretable, knowledge-driven paradigm that transcends purely data-driven approaches. However, we find that plausibility is not prediction: despite producing biologically plausible explanations, these methods fail to capture perturbation-specific effects: systematically overestimating differential expression, often underperforming a simple gene-frequency baseline in aggregate evaluations, and collapsing to chance-level performance at the per-gene level. This reveals a reliance on intrinsic gene response tendencies rather than true perturbation reasoning. We trace this failure to how evidence is presented: existing methods evaluate perturbation-gene pairs in isolation, without exposing how related perturbations differ in their effects on the same gene. To address this limitation, we introduce CORE (Contrastive Organization of Relational Evidence), which reframes prediction as a comparison task by organizing evidence into positive and negative outcomes from related perturbations. Using a biomedical knowledge graph for evidence retrieval, CORE improves calibration and substantially boosts perturbation-specific prediction in both LLM-based and non-LLM settings: for example, on drug-perturbation data, CORE-Reasoning improves Qwen3.5-9B aggregate metrics by up to 28.6%, while on generic perturbation data, CORE-Voting raises macro-per-gene AUROC from chance to 0.703 in average across four cell lines. This highlights contrastive evidence organization as essential to reliable LLM-based perturbation reasoning

基因预测大模型推理对比学习生物信息

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。