arXiv:2602.13791cs.LGcs.AI2026-02被引 1

用共识机制让大模型更准预测基因扰动效应

MechPert: Mechanistic Consensus as an Inductive Bias for Unseen Perturbation Prediction

  • 多智能体生成调控假设,通过共识过滤错误关联
  • 低数据下预测相关性提升10.5%,优于传统相似性方法
  • 适合基因功能研究和实验设计,尤其小样本场景

预测未见基因扰动的转录响应对理解基因调控与规划大规模扰动实验至关重要。现有方法依赖静态知识图谱或基于文本共现的语义相似性提示语言模型,但难以捕捉有向调控逻辑。我们提出 MechPert,一种轻量级框架,引导大模型生成有向调控假设而非仅依赖功能相似性。多个智能体独立提出候选调控因子及置信度;通过共识机制聚合,剔除虚假关联,生成加权邻域用于下游预测。在四个人类细胞系的 Perturb-seq 基准上评估,当仅观测50个扰动(N=50)时,MechPert相较基于相似性的基线,皮尔逊相关性最高提升10.5%。在实验设计中,MechPert选取的锚基因在已知通路细胞系中表现优于标准网络中心性策略,最高提升46%。

原文摘要 · Abstract (English)

Predicting transcriptional responses to unseen genetic perturbations is essential for understanding gene regulation and prioritizing large-scale perturbation experiments. Existing approaches either rely on static, potentially incomplete knowledge graphs, or prompt language models for functionally similar genes, retrieving associations shaped by symmetric co-occurrence in scientific text rather than directed regulatory logic. We introduce MechPert, a lightweight framework that encourages LLM agents to generate directed regulatory hypotheses rather than relying solely on functional similarity. Multiple agents independently propose candidate regulators with associated confidence scores; these are aggregated through a consensus mechanism that filters spurious associations, producing weighted neighborhoods for downstream prediction. We evaluate MechPert on Perturb-seq benchmarks across four human cell lines. For perturbation prediction in low-data regimes ($N=50$ observed perturbations), MechPert improves Pearson correlation by up to 10.5\% over similarity-based baselines. For experimental design, MechPert-selected anchor genes outperform standard network centrality heuristics by up to 46\% in well-characterized cell lines.

基因调控大模型应用实验设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。