用多智能体协作自动预测蛋白质突变,提升实验成功率。
Rank-and-Reason: Multi-Agent Collaboration Accelerates Zero-Shot Protein Mutation Prediction
- 两阶段框架:先排序候选突变,再用推理审计结构合理性。
- 在ProteinGym上相关性达0.551,顶级命中率提升367%。
- 适合低资源蛋白工程,减少人工依赖,加速实验验证。
零样本突变预测对低资源蛋白质工程至关重要,但现有蛋白质语言模型(PLMs)常产生看似可信却违背基本生物物理规律的结果。当前候选突变的选择依赖人工专家审核,效率低且主观性强。为此,我们提出Rank-and-Reason(VenusRAR)框架,通过两阶段多智能体协作自动化该流程,最大化湿实验预期适应度。在排序阶段,计算专家与虚拟生物学家构建上下文感知的多模态集成模型,在ProteinGym上实现0.551的斯皮尔曼相关性(优于原有0.518)。在推理阶段,智能体小组采用思维链推理,依据几何与结构约束审计候选突变,在ProteinGym-DMS99上使前5名命中率最高提升367%。对Cas12i3核酸酶的湿实验验证进一步证明其有效性,阳性率达46.7%,并发现两个新突变分别实现4.23倍和5.05倍活性提升。代码与数据集已开源于GitHub。
原文摘要 · Abstract (English)
Zero-shot mutation prediction is vital for low-resource protein engineering, yet existing protein language models (PLMs) often yield statistically confident results that ignore fundamental biophysical constraints. Currently, selecting candidates for wet-lab validation relies on manual expert auditing of PLM outputs, a process that is inefficient, subjective, and highly dependent on domain expertise. To address this, we propose Rank-and-Reason (VenusRAR), a two-stage agentic framework to automate this workflow and maximize expected wet-lab fitness. In the Rank-Stage, a Computational Expert and Virtual Biologist aggregate a context-aware multi-modal ensemble, establishing a new Spearman correlation record of 0.551 (vs. 0.518) on ProteinGym. In the Reason-Stage, an agentic Expert Panel employs chain-of-thought reasoning to audit candidates against geometric and structural constraints, improving the Top-5 Hit Rate by up to 367% on ProteinGym-DMS99. The wet-lab validation on Cas12i3 nuclease further confirms the framework's efficacy, achieving a 46.7% positive rate and identifying two novel mutants with 4.23-fold and 5.05-fold activity improvements. Code and datasets are released on GitHub (https://github.com/ai4protein/VenusRAR/).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。