用大模型和知识图谱加速抗生素发现,避免重复研发。
Accelerating Antibiotic Discovery with Large Language Models and Knowledge Graphs
- 构建知识图谱融合微生物与化学文献,自动识别已知活性化合物
- 在73个候选菌株中发现12个无抗生素活性的负面案例
- 可公开共享知识图谱与交互界面,助力科研人员高效筛选
抗生素新药发现对应对日益严重的抗菌药物耐药性(AMR)至关重要。然而制药行业面临成本超10亿美元、周期长且失败率高,还常重复发现已知化合物。我们提出一种基于大语言模型的流水线,作为预警系统,检测已有抗生素活性证据,防止昂贵的重复研究。该系统将生物体与化学文献整合为知识图谱(KG),实现分类学精准定位、同义词处理及多层级证据分类。我们在一个包含73个潜在产抗生素微生物的私有列表上测试了该流程,揭示了12个负面结果用于评估。结果表明,该流程在证据审查中有效减少假阴性,加速决策。负面案例的知识图谱及交互式用户界面将公开发布。
原文摘要 · Abstract (English)
The discovery of novel antibiotics is critical to address the growing antimicrobial resistance (AMR). However, pharmaceutical industries face high costs (over $1 billion), long timelines, and a high failure rate, worsened by the rediscovery of known compounds. We propose an LLM-based pipeline that acts as an alarm system, detecting prior evidence of antibiotic activity to prevent costly rediscoveries. The system integrates organism and chemical literature into a Knowledge Graph (KG), ensuring taxonomic resolution, synonym handling, and multi-level evidence classification. We tested the pipeline on a private list of 73 potential antibiotic-producing organisms, disclosing 12 negative hits for evaluation. The results highlight the effectiveness of the pipeline for evidence reviewing, reducing false negatives, and accelerating decision-making. The KG for negative hits and the user interface for interactive exploration will be made publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。