arXiv:2505.24584cs.LGcs.AI2025-05被引 7

用AI自动生成化工放大所需的工程图纸,打通研发到生产的最后一环。

AutoChemSchematic AI: Agentic Physics-Aware Automation for Chemical Manufacturing Scale-Up

  • 用专用小模型+知识图谱+仿真器闭环生成可工业化的流程图和仪表图。
  • 生成的工艺方案经仿真验证,对1020种化学品实现高保真流程设计。
  • 通过结构剪枝和高效推理技术,让大模型在本地运行更快速、低资源。

生成式AI加速了新化学品与材料的发现,但将其放大至工业生产仍面临合成鸿沟——需重新开发制造工艺。这需要详细的工程蓝图:用于设备布局和物料/能量流的PFD,以及用于工厂操作的PID。当前AI尚无法可靠生成这些关键图纸,成为规模化生产的根本障碍。本文提出一个闭环保理、物理感知的自动化框架,用于生成符合工业标准的PFD与PID。该框架集成三项核心技术:(1) 针对PFD/PID自动生成训练的领域专用小语言模型(SLMs);(2) 包含1020余种化学品工艺与仪表描述的分层知识图谱,支持图检索增强生成(GRAG);(3) 开源化学过程仿真器,用于建模、模拟、优化与分析新工艺。SLMs通过多阶段合成数据集训练,并结合仿真器闭环验证可行性。为提升效率,框架采用基于重要性启发的结构剪枝(宽度与深度),减少模型规模同时保持精度,并引入FlashAttention、Lookahead Decoding、PagedAttention与KV缓存量化等先进推理优化技术。实验表明,本框架生成的工艺描述具有高保真度且经仿真验证可行。

原文摘要 · Abstract (English)

Recent advances in generative AI have accelerated the discovery of novel chemicals and materials. However, scaling these discoveries to industrial production remains a major bottleneck due to the synthesis gap -- the need to develop entirely new manufacturing processes. This challenge requires detailed engineering blueprints: PFDs for equipment layouts and material/energy flows, and PIDs for process plant operations. Current AI systems cannot yet reliably generate these critical engineering schematics, creating a fundamental obstacle to manufacturing scale-up of novel discoveries. We present a closed-loop, physics-aware framework for automated generation of industrially viable PFDs and PIDs. The framework integrates three key components: (1) domain-specialized small language models (SLMs) trained for auto-generation of PFDs and PIDs, (2) a hierarchical knowledge graph containing process flow and instrumentation descriptions for 1,020+ chemicals for Graph Retrieval-Augmented Generation (GRAG), and (3) an open-source chemical process simulator for modeling, simulation, optimization, and analysis of novel chemical processes. The SLMs are trained through a multi-stage pipeline on synthetic datasets, with process simulator-in-the-loop validation ensuring feasibility. To enhance computational efficiency, the framework implements structural pruning (width and depth) guided by importance heuristics to reduce language model size while preserving accuracy, followed by advanced inference optimizations including FlashAttention, Lookahead Decoding, PagedAttention with KV-cache quantization, and Test-Time Inference Scaling. Experimental results demonstrate that our framework generates simulator-validated process descriptions with high fidelity.

化工自动化生成式AI工艺设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。