用AI从化工文献中自动提取催化反应数据并支持自然语言分析
AgentCAT: An LLM Agent for Extracting and Analyzing Catalytic Reaction Data from Chemical Engineering Literature
- 基于模式演进的提取管道,稳定抓取复杂反应数据
- 构建依赖感知的知识图谱,关联催化剂与反应证据
- 支持自然语言查询,适合研究人员跨论文分析
本文提出一种名为AgentCAT的大语言模型代理,可从化学工程文献中提取并分析催化反应数据,并支持基于自然语言的交互式数据分析。该方法旨在解决化学工程领域长期存在的数据瓶颈问题,其自然语言交互功能对研究社区友好。AgentCAT对催化反应数据提取任务进行了形式化抽象,以人工智能友好的方式呈现挑战,有助于吸引更多关注。由于催化过程复杂,反应数据在基元步骤、分子行为、测量证据等方面存在复杂依赖关系,导致提取正确性和完整性难以保证。AgentCAT通过四项技术贡献应对这一挑战:(1) 基于模式演进的提取流水线,实现对化工文献的稳健数据抽取;(2) 依赖感知的反应网络知识图谱,连接催化剂/活性位点、合成衍生描述符、机理主张与证据、宏观结果,保持过程耦合与可追溯性;(3) 通用查询模块,支持对构建图谱的自然语言探索与可视化,实现跨论文分析;(4) 在约800篇同行评审化学工程出版物上的评估,验证了AgentCAT的有效性。
原文摘要 · Abstract (English)
This paper presents a large language model (LLM) agent named AgentCAT, which extracts and analyzes catalytic reaction data from chemical engineering papers, %and supports natural language based interactive analysis of the extracted data. AgentCAT serves as an alternative to overcome the long-standing data bottleneck in chemical engineering field, and its natural language based interactive data analysis functionality is friendly to the community. AgentCAT also presents a formal abstraction and challenge analysis of the catalytic reaction data extraction task in an artificial intelligence-friendly manner. This abstraction would help the artificial intelligence community understand this problem and in turn would attract more attention to address it. Technically, the complex catalytic process leads to complicated dependency structure in catalytic reaction data with respect to elementary reaction steps, molecular behaviors, measurement evidence, etc. This dependency structure makes it challenging to guarantee the correctness and completeness of data extraction, as well as representing them for analysis. AgentCAT addresses this challenge and it makes four folds of technical contributions: (1) a schema-governed extraction pipeline with progressive schema evolution, enabling robust data extraction from chemical engineering papers; (2) a dependency-aware reaction-network knowledge graph that links catalysts/active sites, synthesis-derived descriptors, mechanistic claims with evidence, and macroscopic outcomes, preserving process coupling and traceability; (3) a general querying module that supports natural-language exploration and visualization over the constructed graph for cross-paper analysis; (4) an evaluation on $\sim$800 peer-reviewed chemical engineering publications demonstrating the effectiveness of AgentCAT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。