arXiv:2410.21312cs.LGcs.AI2024-10被引 5

首个智能专利分析代理,自动解析药企专利并提取关键化学结构。

$\texttt{PatentAgent}$: Intelligent Agent for Automated Pharmaceutical Patent Analysis

  • 用大模型构建三模块一体化系统,覆盖问答、图像转分子式、核心结构识别
  • 图像转分子结构准确率提升2.46%~8.37%,核心结构识别提升7.15%~7.62%
  • 适合药物研发人员快速挖掘专利中的化学信息,提升创新效率

药品专利在生物化学产业中至关重要,尤其在新药研发阶段,为研究人员提供早期数据、实验结果与研究洞见。随着机器学习的发展,专利分析已从人工转向自动化工具辅助。然而,目前仍缺乏能覆盖专利阅读、核心化学结构识别等全流程的统一智能代理。本文提出首个该领域的智能代理——PatentAgent,利用大语言模型理解指令与任务需求,包含三个端到端模块:PA-QA(专利问答)、PA-Img2Mol(图像转分子结构)和PA-CoreId(核心化学结构识别),满足科研人员在专利分析中的关键需求。各模块在更新算法与框架协同设计下表现优异:PA-Img2Mol 在 CLEF、JPO、UOB 与 USPTO 基准上准确率提升 2.46% 至 8.37%;PA-CoreId 在 PatentNetML 基准上准确率提升 7.15% 至 7.62%。代码与数据集将公开。

原文摘要 · Abstract (English)

Pharmaceutical patents play a vital role in biochemical industries, especially in drug discovery, providing researchers with unique early access to data, experimental results, and research insights. With the advancement of machine learning, patent analysis has evolved from manual labor to tasks assisted by automatic tools. However, there still lacks an unified agent that assists every aspect of patent analysis, from patent reading to core chemical identification. Leveraging the capabilities of Large Language Models (LLMs) to understand requests and follow instructions, we introduce the $\textbf{first}$ intelligent agent in this domain, $\texttt{PatentAgent}$, poised to advance and potentially revolutionize the landscape of pharmaceutical research. $\texttt{PatentAgent}$ comprises three key end-to-end modules -- $\textit{PA-QA}$, $\textit{PA-Img2Mol}$, and $\textit{PA-CoreId}$ -- that respectively perform (1) patent question-answering, (2) image-to-molecular-structure conversion, and (3) core chemical structure identification, addressing the essential needs of scientists and practitioners in pharmaceutical patent analysis. Each module of $\texttt{PatentAgent}$ demonstrates significant effectiveness with the updated algorithm and the synergistic design of $\texttt{PatentAgent}$ framework. $\textit{PA-Img2Mol}$ outperforms existing methods across CLEF, JPO, UOB, and USPTO patent benchmarks with an accuracy gain between 2.46% and 8.37% while $\textit{PA-CoreId}$ realizes accuracy improvement ranging from 7.15% to 7.62% on PatentNetML benchmark. Our code and dataset will be publicly available.

专利分析药物研发大模型应用分子生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。