arXiv:2502.21290cs.AIcs.LG2025-02ICLR被引 21

用大模型理解生物扰动实验,提升数据分析效率与可解释性

Contextualizing biological perturbation experiments through language

  • 构建新基准PerturbQA,聚焦未见扰动的表达预测与通路变化推理
  • 现有方法在新任务上表现差,准确率不足40%
  • 提出Summer框架,结合检索与摘要,效果优于当前最优模型

高内涵扰动实验可前所未有的分辨率研究生物分子系统,但实验与分析成本限制了广泛应用。机器学习有望高效探索扰动空间并从中提取新见解,但现有方法忽视生物学语义丰富性,目标与下游分析脱节。本文假设大语言模型(LLMs)是表征复杂生物关系和解释实验结果的自然媒介。提出PerturbQA基准,用于结构化推理扰动实验,不同于仅检验已有知识的现有基准,其聚焦开放问题:未见扰动的差异表达预测、方向变化预测及基因集富集分析。评估了主流机器学习与统计方法,以及标准LLM推理策略,发现当前方法在PerturbQA上表现不佳。作为可行性验证,提出Summer(SUMMarize, retrievE, and answeR)框架,一种简单且领域感知的LLM方案,在多个任务上达到或超越当前最佳水平。代码与数据已公开于https://github.com/genentech/PerturbQA。

原文摘要 · Abstract (English)

High-content perturbation experiments allow scientists to probe biomolecular systems at unprecedented resolution, but experimental and analysis costs pose significant barriers to widespread adoption. Machine learning has the potential to guide efficient exploration of the perturbation space and extract novel insights from these data. However, current approaches neglect the semantic richness of the relevant biology, and their objectives are misaligned with downstream biological analyses. In this paper, we hypothesize that large language models (LLMs) present a natural medium for representing complex biological relationships and rationalizing experimental outcomes. We propose PerturbQA, a benchmark for structured reasoning over perturbation experiments. Unlike current benchmarks that primarily interrogate existing knowledge, PerturbQA is inspired by open problems in perturbation modeling: prediction of differential expression and change of direction for unseen perturbations, and gene set enrichment. We evaluate state-of-the-art machine learning and statistical approaches for modeling perturbations, as well as standard LLM reasoning strategies, and we find that current methods perform poorly on PerturbQA. As a proof of feasibility, we introduce Summer (SUMMarize, retrievE, and answeR, a simple, domain-informed LLM framework that matches or exceeds the current state-of-the-art. Our code and data are publicly available at https://github.com/genentech/PerturbQA.

生物信息大模型扰动分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。