arXiv:2608.16419cs.LGcs.AI2026-08

用基因扰动数据训练大模型,让其自动学会生物推理。

PertMind: Eliciting Emergent Biological Reasoning in LLM via Reinforcement Learning on Cellular Perturbation Data

  • 以基因扰动响应作为奖励信号,通过强化学习训练模型。
  • 在未见细胞环境中提升基因响应预测准确率,且无需额外微调。
  • 适用于反向扰动识别、多扰动推理等复杂任务,适合生物信息研究者。

大语言模型虽能描述机制,但可扩展的后训练仍依赖昂贵的人工标注生物推理轨迹。本文提出利用细胞扰动图谱作为强化学习环境,通过测量的基因响应提供可计算的奖励信号。我们引入PertMind,结合可信轨迹监督初始化与基因、通路及格式层面的强化信号。仅在正向扰动-响应预测任务上训练,PertMind在未见细胞背景下提升了响应推断能力,同时保持通用语言能力。该模型无需特定任务微调即可迁移至反向扰动识别、双扰动推理、表型筛选优先级排序及生物过程解释。PertMind还生成了在多尺度下游任务中表现优异的基因、细胞和供体表示。结果支持假设:基于实验终点的强化学习可聚焦预训练模型已具备的可复用生物策略。更广泛地,扰动衍生的强化学习为将不断扩大的实验图谱转化为通用生物推理训练环境提供了可扩展路径。

原文摘要 · Abstract (English)

Large language models can describe mechanisms, yet scalable post-training still depends on costly, manually curated biological reasoning traces. Here we show that cellular perturbation atlases can instead become reinforcement-learning environments, where measured gene responses provide computable rewards for biological reasoning. We introduce PertMind, which combines trusted-trajectory supervised initialization with gene-, pathway-, and format-level reinforcement signals. Trained only on forward perturbation-response prediction, PertMind improved response inference in unseen cellular contexts while retaining general language capabilities. It also transferred without task-specific post-training to reverse perturbation identification, double-perturbation reasoning, phenotypic-screen prioritization, and biological-process interpretation. PertMind further generated biological profiles that supported competitive gene, cell, and donor representations across multiscale downstream tasks. These results support the hypothesis that reinforcement on experimental endpoints can concentrate reusable biological strategies already accessible to pretrained models. More broadly, perturbation-derived reinforcement learning offers a scalable route for transforming expanding experimental atlases into training environments for general-purpose biological reasoning.

生物推理强化学习基因扰动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。