用生成式AI自动搜寻最佳生物标志物组合,省去大量实验人力。
Revolutionizing Biomarker Discovery: Leveraging Generative AI for Bio-Knowledge-Embedded Continuous Space Exploration
- 构建多智能体系统自动生成标志物组合与预测精度数据
- 在三个真实数据集上实现高效且稳定的标志物筛选
- 适合需要快速发现可靠生物标志物的生物医药研究者
生物标志物发现对推动个性化医疗至关重要,有助于疾病诊断、预后判断和疗效评估。传统方法依赖大量实验和统计分析,耗时长、需深厚领域知识,且受限于生物系统的复杂性。为解决此问题,我们提出一种新框架,旨在无需大量人工干预即可自动识别有效生物标志物子集。受生成式AI成功启发,我们将生物标志物识别的复杂知识压缩至连续嵌入空间,以提升搜索效率。该框架包含两个核心模块:1)训练数据准备,利用多智能体系统自动收集标志物子集及其对应预测准确率的数据对;2)嵌入优化生成,采用编码器-评估器-解码器学习范式,将数据知识压缩到连续空间,并结合基于梯度的搜索与自回归重构技术,高效寻找最优标志物子集。我们在三个真实世界数据集上进行了广泛实验,验证了该方法在效率、鲁棒性和有效性方面的优势。
原文摘要 · Abstract (English)
Biomarker discovery is vital in advancing personalized medicine, offering insights into disease diagnosis, prognosis, and therapeutic efficacy. Traditionally, the identification and validation of biomarkers heavily depend on extensive experiments and statistical analyses. These approaches are time-consuming, demand extensive domain expertise, and are constrained by the complexity of biological systems. These limitations motivate us to ask: Can we automatically identify the effective biomarker subset without substantial human efforts? Inspired by the success of generative AI, we think that the intricate knowledge of biomarker identification can be compressed into a continuous embedding space, thus enhancing the search for better biomarkers. Thus, we propose a new biomarker identification framework with two important modules:1) training data preparation and 2) embedding-optimization-generation. The first module uses a multi-agent system to automatically collect pairs of biomarker subsets and their corresponding prediction accuracy as training data. These data establish a strong knowledge base for biomarker identification. The second module employs an encoder-evaluator-decoder learning paradigm to compress the knowledge of the collected data into a continuous space. Then, it utilizes gradient-based search techniques and autoregressive-based reconstruction to efficiently identify the optimal subset of biomarkers. Finally, we conduct extensive experiments on three real-world datasets to show the efficiency, robustness, and effectiveness of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。