用合成推理轨迹提升大模型预测基因扰动效果
SynthPert: Enhancing LLM Biological Reasoning via Synthetic Reasoning Traces for Cellular Perturbation Prediction
- 用前沿模型生成合成推理路径,监督微调大模型
- 在未见RPE1细胞上达87%准确率,超越原始模型
- 仅需2%高质量数据即实现性能提升,适合生物推理
预测基因扰动对细胞的响应是系统生物学中的核心挑战,对药物发现和虚拟细胞建模至关重要。尽管大语言模型(LLMs)在生物推理中展现潜力,但其在扰动预测中的应用仍受限于难以适配结构化实验数据。我们提出SynthPert,通过前沿模型生成的合成推理轨迹对LLM进行有监督微调,显著提升性能。在PerturbQA基准测试中,该方法不仅达到当前最优水平,还超越了生成训练数据的前沿模型。结果揭示三个关键发现:(1) 即使部分不准确,合成推理轨迹仍能有效提炼生物知识;(2) 实现跨细胞类型泛化,在未见RPE1细胞上达87%准确率;(3) 仅使用2%经过质量筛选的训练数据,性能提升依然显著。本工作证明了合成推理蒸馏在增强领域特定推理能力方面的有效性。
原文摘要 · Abstract (English)
Predicting cellular responses to genetic perturbations represents a fundamental challenge in systems biology, critical for advancing therapeutic discovery and virtual cell modeling. While large language models (LLMs) show promise for biological reasoning, their application to perturbation prediction remains underexplored due to challenges in adapting them to structured experimental data. We present SynthPert, a novel method that enhances LLM performance through supervised fine-tuning on synthetic reasoning traces generated by frontier models. Using the PerturbQA benchmark, we demonstrate that our approach not only achieves state-of-the-art performance but surpasses the capabilities of the frontier model that generated the training data. Our results reveal three key insights: (1) Synthetic reasoning traces effectively distill biological knowledge even when partially inaccurate, (2) This approach enables cross-cell-type generalization with 87% accuracy on unseen RPE1 cells, and (3) Performance gains persist despite using only 2% of quality-filtered training data. This work shows the effectiveness of synthetic reasoning distillation for enhancing domain-specific reasoning in LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。