arXiv:2504.20598cs.IR2025-04被引 1

用NLP从专利中提取药物制造信息,填补行业数据空白

Natural Language Processing tools for Pharmaceutical Manufacturing Information Extraction from Patents

  • 用主题模型与聚类定位制造相关文本段落
  • 实体识别模型F1值达84.2%,性能可比肩同类工作
  • 适合制药数据工程、AI辅助研发人员参考

过去几十年,药品制造及其他生命周期环节的数据大量公开,但多数信息未结构化,难以用于机器学习。为此,自然语言处理工具已在生物医学和化学领域用于构建数据库,推动了药物发现与治疗的智能化。然而,在药物制剂制造领域,现有研究和数据集仍显不足。本文旨在探索并适配其他领域的NLP技术,从专利中提取原辅料、工艺操作及条件等制造信息,涵盖初级与次级制造流程。为此开发了两个互补模型:首个模型通过无监督方法(结合隐含狄利克雷分布与k均值聚类)识别含制造数据的文本片段,其预测结果与人工标注的科恩κ系数高于90%;第二个模型为命名实体识别系统,基于深度神经网络,微平均F1得分为84.2%,与现有工作相当。文中还讨论了该工具在实际数据抽取中的适用性与改进方向。

原文摘要 · Abstract (English)

Abundant and diverse data on medicines manufacturing and other lifecycle components has been made easily accessible in the last decades. However, a significant proportion of this information is characterised by not being tabulated and usable for machine learning purposes. Thus, natural language processing tools have been used to build databases in domains such as biomedical and chemical to address this limitation. This has allowed the development of artificial intelligence applications, which have improved drug discovery and treatments. In the pharmaceutical manufacturing context, some initiatives and datasets for primary processing can be found, but the manufacturing of drug products is an area which is still lacking, to the best of our knowledge. This works aims to explore and adapt NLP tools used in other domains to extract information on both primary and secondary manufacturing, employing patents as the main source of data. Thus, two independent, but complementary, models were developed comprising a method to select fragments of text that contain manufacturing data, and a named entity recognition system that enables extracting information on operations, materials, and conditions of a process. For the first model, the identification of relevant sections was achieved using an unsupervised approach combining Latent Dirichlet Allocation and k-Means clustering. The performance of this model measured as a Cohen's kappa between model output and manual revision was higher than 90%. NER model consisted of a deep neural network, and an f1-score micro average of 84.2% was obtained which is comparable to other works. Some considerations for these tools to be used in data extraction are discussed throughout this document.

NLP制药制造信息提取专利分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。