AI平台自主训练蛋白互作模型并提炼可解释规则,提升预测准确率与可理解性。
Agentic AI platforms for autonomous training and rule induction of human-human and virus-human protein-protein interactions
- 构建双平台:一自动训练预测模型,一生成人类与病毒蛋白互作用的可读规则。
- 人类间与人-病毒互作预测准确率分别达87.3%和86.5%,基于三重蛋白不相交交叉验证。
- 规则来自蛋白嵌入等多源特征,与模型解释结果一致,适合生物医学研究者使用。
我们指导一个AI代理构建两个独立的智能体平台:一个用于自主训练预测人类间及人-病毒蛋白-蛋白互作(PPI)的机器学习模型,另一个用于推导描述此类互作的显式通用规则。首个平台由五个智能体组成,分别负责自主数据采集、数据验证、特征嵌入、模型设计以及在三重蛋白不相交交叉验证数据集上的训练与验证。对于人类间和人-病毒PPI,最终的三重蛋白不相交集成模型分别达到87.3%和86.5%的准确率。为保证可解释性,第二个平台用从蛋白嵌入、理化自协方差描述符、细胞区室注释、通路域重叠及图上下文提取的人类可读规则替代机器学习预测。人类间互作用由两条规则定义,而人-病毒互作用则由一组加权规则构成。这些规则与第一个平台中基于SHAP识别的关键特征高度一致。本工作展示了智能体在从数据规划到执行、从规则推导到解释的全流程自主能力,为多种应用打开新可能。
原文摘要 · Abstract (English)
We instruct an AI agent to construct two separate agentic AI platforms: one for autonomous training of predictive ML models for human-human and virus-human PPI, and the other for inducing explicit general rules governing human-human and virus-human PPI. The first agentic AI platform for autonomous training of predictive ML models for PPI is designed to consist of five AI agents that handle autonomous data collection, data verification, feature embedding, model design, and training and validation on three-way protein-disjoint cross-fold datasets. For human-human and human-virus PPIs, the final three-way protein-disjoint ensemble achieves an accuracy of 87.3% and 86.5%, respectively. For cross-checking and interpretability purposes, the second agentic AI platform is designed to replace ML predictions with human-readable rules derived from protein embeddings, physicochemical autocovariance descriptors, compartment annotations, pathway-domain overlap, and graph contexts. For human-human PPI, it is defined by a two-rule induction, whereas human-virus is induced by a more complex set of weighted rules. The rules induced by the second agentic platform align with the SHAP-identified features from the predictive ML models built by the first agentic platform. Taken together, our work demonstrates the agentic AI's ability to orchestrate from data planning to execution, and from rule induction to explanation in ML, opening the door to various applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。