arXiv:2506.07591cs.AIq-bio.QM2025-06被引 2

用语言模型自动从多组学数据生成可验证的科学假说。

Automating Exploratory Multiomics Research via Language Models

  • 分模块模拟科研流程,用统一图结构管理生物实体关系。
  • 在10个临床多组学数据集上生成360个假说,经外部数据验证。
  • 适合需要高效探索性研究的生物医学领域研究人员。

本文提出PROTEUS,一个全自动系统,可从原始数据文件生成数据驱动的假说。针对临床蛋白质基因组学这一关键领域,该系统通过分离模块模拟科学过程的不同阶段,涵盖开放式数据探索、特定统计分析及假说生成。它以生物实体间关系的形式表达研究方向、工具与结果,并利用统一图结构管理复杂研究流程。我们将PROTEUS应用于10个已发表研究的临床多组学数据集,共生成360个假说。通过外部数据验证和自动开放式评分进行评估。经过探索性与迭代式研究,系统能有效处理高通量异构多组学数据,生成兼具可靠性与新颖性的假说。除加速多组学分析外,PROTEUS还为通用自主系统向专业科学领域定制提供了路径,实现从数据出发的开放式假说生成。

原文摘要 · Abstract (English)

This paper introduces PROTEUS, a fully automated system that produces data-driven hypotheses from raw data files. We apply PROTEUS to clinical proteogenomics, a field where effective downstream data analysis and hypothesis proposal is crucial for producing novel discoveries. PROTEUS uses separate modules to simulate different stages of the scientific process, from open-ended data exploration to specific statistical analysis and hypothesis proposal. It formulates research directions, tools, and results in terms of relationships between biological entities, using unified graph structures to manage complex research processes. We applied PROTEUS to 10 clinical multiomics datasets from published research, arriving at 360 total hypotheses. Results were evaluated through external data validation and automatic open-ended scoring. Through exploratory and iterative research, the system can navigate high-throughput and heterogeneous multiomics data to arrive at hypotheses that balance reliability and novelty. In addition to accelerating multiomic analysis, PROTEUS represents a path towards tailoring general autonomous systems to specialized scientific domains to achieve open-ended hypothesis generation from data.

多组学假说生成自动化科研

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。