用大模型自动分析蛋白组数据,自动生成可验证的科学假说。
Automating Exploratory Proteomics Research via Language Models
- 基于大模型分层规划,自动调用生物信息工具迭代优化分析流程。
- 在12个数据集上生成191个假说,经专家评审与自动评分均表现良好。
- 适合希望高效挖掘蛋白组数据的科研人员,尤其擅长多类型样本分析。
随着人工智能发展,其在科学中的角色正从模拟复杂问题转向自动化完整研究流程并催生新发现。这需要基于真实科学数据的专用通用模型,以及模仿人类科研方法的迭代探索框架。本文提出PROTEUS,一个从原始蛋白组数据出发的全自动科学发现系统。PROTEUS利用大语言模型(LLMs)进行分层规划,执行专业生物信息工具,并迭代优化分析工作流,生成高质量科学假说。系统输入蛋白组数据后,无需人工干预即可输出完整研究目标、分析结果和新颖生物学假说。我们在12个来自不同生物样本(如免疫细胞、肿瘤)和样本类型(单细胞与批量)的蛋白组数据集上评估PROTEUS,生成191个科学假说。通过5项指标的自动大模型评分及专家详细评审,结果表明PROTEUS持续产出可靠、逻辑连贯且与现有文献高度一致的结果,同时提出可评估的新颖假说。其灵活架构支持多种分析工具无缝集成,并适应不同蛋白组数据类型。通过自动化复杂蛋白组分析流程与假说生成,PROTEUS有望显著加速蛋白组学研究进程,助力研究人员高效探索大规模数据集并揭示生物学洞见。
原文摘要 · Abstract (English)
With the development of artificial intelligence, its contribution to science is evolving from simulating a complex problem to automating entire research processes and producing novel discoveries. Achieving this advancement requires both specialized general models grounded in real-world scientific data and iterative, exploratory frameworks that mirror human scientific methodologies. In this paper, we present PROTEUS, a fully automated system for scientific discovery from raw proteomics data. PROTEUS uses large language models (LLMs) to perform hierarchical planning, execute specialized bioinformatics tools, and iteratively refine analysis workflows to generate high-quality scientific hypotheses. The system takes proteomics datasets as input and produces a comprehensive set of research objectives, analysis results, and novel biological hypotheses without human intervention. We evaluated PROTEUS on 12 proteomics datasets collected from various biological samples (e.g. immune cells, tumors) and different sample types (single-cell and bulk), generating 191 scientific hypotheses. These were assessed using both automatic LLM-based scoring on 5 metrics and detailed reviews from human experts. Results demonstrate that PROTEUS consistently produces reliable, logically coherent results that align well with existing literature while also proposing novel, evaluable hypotheses. The system's flexible architecture facilitates seamless integration of diverse analysis tools and adaptation to different proteomics data types. By automating complex proteomics analysis workflows and hypothesis generation, PROTEUS has the potential to considerably accelerate the pace of scientific discovery in proteomics research, enabling researchers to efficiently explore large-scale datasets and uncover biological insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。