arXiv:2608.05120cs.LGcs.CE2026-08

用大模型指导符号回归,加速化学反应动力学模型发现。

DASyR-LLM: Domain-Aware Symbolic Regression with LLMs for Kinetic Model Discovery

论文配图:DASyR-LLM: Domain-Aware Symbolic Regression with LLMs for Kinetic Model Discovery
图 1 · 摘自论文原文
  • 将大模型嵌入迭代符号回归流程,实时提供化学合理性评估与新模型建议。
  • 相比顶尖方法减少41.7%~79.3%的迭代次数,超半数实验直接命中正确模型结构。
  • 适合需要快速构建可解释反应模型的化工与生物过程研究者。

化学工程中的动力学模型发现是一项核心挑战,准确的速率表达式对理解与控制化学及生物过程至关重要。符号回归(SR)作为一种数据驱动的方法,能发现可解释的动力学模型,但通常不引入领域知识,常探索出物理化学上不合理的形式。大语言模型(LLM)为注入领域专长提供了新路径。本文提出一种基于LLM引导的符号回归框架,将LLM模块嵌入迭代式符号回归算法中,实现自动化动力学模型发现。在每轮迭代中,LLM承担双重角色:(1)对最优候选模型进行定性物化化学批判;(2)基于已有模型与内嵌化学知识生成新候选速率表达式。该框架在四个复杂度递增的模拟案例中验证,涵盖异相催化与生物过程系统。结果表明,相较当前最先进方法,该框架使识别真实模型的迭代次数减少41.7%至79.3%,且在超过一半的运行中,LLM直接提出了正确的模型结构。在实际场景中,每次迭代通常需一次湿实验,此举显著降低实验工作量。独立验证集上的预测性能相当,所有案例$R^2>0.98$。消融分析显示,SR组件与LLM均贡献性能,小规模LLM仍保持较高发现效率。研究证明,LLM可有效将领域知识融入科学模型发现,为实现全自动、领域感知的动力学建模流程铺平道路。

原文摘要 · Abstract (English)

Kinetic model discovery is a central challenge in chemical engineering, as accurate rate expressions are essential for understanding and controlling chemical and biological processes. Symbolic regression (SR) has emerged as a powerful data-driven approach for identifying interpretable kinetic models, but usually operates without domain knowledge, often exploring physicochemically implausible models. Large language models (LLMs) offer a promising avenue for injecting domain expertise into this search. Here, we introduce an LLM-guided SR framework, embedding an LLM module within an iterative SR algorithm for automated kinetic model discovery. The LLM performs two roles at each iteration: (1) a qualitative physicochemical critique of the best SR candidates, and (2) the proposal of new candidate rate expressions guided by the SR-generated models and embedded chemical knowledge. Our framework is evaluated on four in silico case studies of increasing complexity, spanning heterogeneous catalysis and bioprocess systems. Results show the LLM-guided framework reduces iterations to identify the ground-truth model by $41.7-79.3\%$ versus a state-of-the-art SR framework, with the LLM directly proposing the correct model structure in over half of the guided runs. In practical settings, where each iteration typically requires a new wet-lab experiment, this translates into a substantial reduction in experimental effort. Predictive performance on an independent validation set is equivalent between both approaches, with $R^2>0.98$ in all case studies. Ablation studies indicate that both the SR component and the LLM scale contribute to this performance, with a reduced-size LLM largely retaining discovery efficiency. These findings demonstrate that LLMs can effectively inject domain knowledge into scientific model discovery, paving the way toward fully automated, domain-aware kinetic modelling pipelines.

符号回归大模型动力学建模化学工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。