用大模型零样本推断生物因果关系,效果堪比专业研究。
Large Language Models for Zero-shot Inference of Causal Structures in Biology
- 通过提示工程与文献检索增强,让小模型也能捕捉生物网络因果结构。
- 在上百变量、数千假设上验证,模型推理与真实干预数据高度一致。
- 适合生物发现、因果推理初学者或快速生成假说的研究者使用。
基因、蛋白质等生物实体通过复杂的因果分子网络相互影响。这些网络中的因果关系受潜在变量调节,且常具有细胞环境特异性,实际刻画极具挑战。本文提出一种新框架,评估大语言模型(LLMs)在生物学中进行零样本因果关系推断的能力。我们基于真实干预数据系统性地检验了来自LLM的因果主张,在超过一百个变量和数千个因果假设上展开验证。同时测试了多种提示策略与检索增强方法,包括大规模且可能存在冲突的科学文献集合。结果表明,经过适当增强与提示设计,即使相对较小的LLM也能捕捉生物系统中重要的因果结构特征。这支持了LLM可作为生物学发现中的知识整合工具,将现有知识转化为可分析的形式。本方法对因果学习、大模型与科学发现交叉领域的诸多问题具有参考价值。
原文摘要 · Abstract (English)
Genes, proteins and other biological entities influence one another via causal molecular networks. Causal relationships in such networks are mediated by complex and diverse mechanisms, through latent variables, and are often specific to cellular context. It remains challenging to characterise such networks in practice. Here, we present a novel framework to evaluate large language models (LLMs) for zero-shot inference of causal relationships in biology. In particular, we systematically evaluate causal claims obtained from an LLM using real-world interventional data. This is done over one hundred variables and thousands of causal hypotheses. Furthermore, we consider several prompting and retrieval-augmentation strategies, including large, and potentially conflicting, collections of scientific articles. Our results show that with tailored augmentation and prompting, even relatively small LLMs can capture meaningful aspects of causal structure in biological systems. This supports the notion that LLMs could act as orchestration tools in biological discovery, by helping to distil current knowledge in ways amenable to downstream analysis. Our approach to assessing LLMs with respect to experimental data is relevant for a broad range of problems at the intersection of causal learning, LLMs and scientific discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。