arXiv:2509.26128cs.AI2025-09被引 1

用大模型从药品说明书构建生物医学知识图谱,覆盖副作用、用法等关键信息。

MEDAKA: Construction of Biomedical Knowledge Graphs Using Large Language Models

  • 用网络爬虫和大模型自动提取说明书文本,构建端到端知识图谱流水线。
  • 创建了包含副作用、禁忌症、用量等10类临床信息的MEDAKA数据集。
  • 适合做药物安全监控与个性化用药推荐,也可拓展至其他领域。

知识图谱(KGs)正被广泛用于以结构化、可解释的形式表示生物医学信息。然而,现有生物医学知识图谱多聚焦于分子相互作用或不良反应,忽视了药品说明书中的丰富数据。本文提出(1)一个可定制的端到端流程,利用网络爬虫与大语言模型从非结构化在线内容中构建知识图谱;(2)通过该方法对公开药品说明书进行处理,构建了一个名为MEDAKA的高质量数据集。该数据集涵盖临床相关属性,包括不良反应、警告、禁忌症、成分、剂量指南、储存说明及物理特性等。我们通过人工检查和基于大模型的评判框架评估其质量,并与现有生物医学知识图谱及数据库对比覆盖范围。预计MEDAKA可支持患者安全监测与药物推荐任务。该流程亦可用于其他领域非结构化文本的知识图谱构建。代码与数据集已公开于 https://github.com/medakakg/medaka。

原文摘要 · Abstract (English)

Knowledge graphs (KGs) are increasingly used to represent biomedical information in structured, interpretable formats. However, existing biomedical KGs often focus narrowly on molecular interactions or adverse events, overlooking the rich data found in drug leaflets. In this work, we present (1) a hackable, end-to-end pipeline to create KGs from unstructured online content using a web scraper and an LLM; and (2) a curated dataset, MEDAKA, generated by applying this method to publicly available drug leaflets. The dataset captures clinically relevant attributes such as side effects, warnings, contraindications, ingredients, dosage guidelines, storage instructions and physical characteristics. We evaluate it through manual inspection and with an LLM-as-a-Judge framework, and compare its coverage with existing biomedical KGs and databases. We expect MEDAKA to support tasks such as patient safety monitoring and drug recommendation. The pipeline can also be used for constructing KGs from unstructured texts in other domains. Code and dataset are available at https://github.com/medakakg/medaka.

知识图谱药物安全大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。