arXiv:2604.00264cs.LG2026-04

用强化学习自动选化学积分器,提速三倍还保准精度。

Autonomous Adaptive Solver Selection for Chemistry Integration via Reinforcement Learning

论文配图:Autonomous Adaptive Solver Selection for Chemistry Integration via Reinforcement Learning
图 1 · 摘自论文原文
  • 把积分器选择建模为马尔可夫决策过程,学全局策略而非只看当下状态。
  • 在0D反应器上平均提速3倍,最高达10.58倍,误差可控且推理开销仅1%。
  • 训练好的策略能直接用于1D火焰模拟,稳定提速2.2倍,适合复杂多物理场仿真。

刚性化学动力学的计算成本仍是反应流模拟的主要瓶颈。传统混合积分策略依赖人工调参或监督预测,只能根据瞬时状态做出局部决策。本文提出一种约束强化学习框架,自主在隐式BDF积分器(CVODE)与准稳态(QSS)求解器间切换。将求解器选择建模为马尔可夫决策过程,智能体学习轨迹感知策略,考虑当前选择对后续误差累积的影响,同时在用户指定精度容忍度下最小化计算开销,通过带在线乘子自适应的拉格朗日奖励实现。在采样的0D均质反应器条件下,强化学习自适应策略实现约3倍平均加速,速度提升范围为1.11×至10.58×,保持106物种正十二烷机制的点火延迟和组分分布准确,并增加约1%推理开销。无需重训练,0D训练策略可迁移至1D逆流扩散火焰,应变率范围为10–2000 s⁻¹,相对CVODE持续获得约2.2倍加速,温度精度接近参考值,且仅在12–15%的空间-时间点选择CVODE。结果表明,该强化学习框架能在满足精度约束前提下学习问题特异性积分策略,为具有空间异质刚性的多物理场系统开启自适应、自优化工作流的新路径。

原文摘要 · Abstract (English)

The computational cost of stiff chemical kinetics remains a dominant bottleneck in reacting-flow simulation, yet hybrid integration strategies are typically driven by hand-tuned heuristics or supervised predictors that make myopic decisions from instantaneous local state. We introduce a constrained reinforcement learning (RL) framework that autonomously selects between an implicit BDF integrator (CVODE) and a quasi-steady-state (QSS) solver during chemistry integration. Solver selection is cast as a Markov decision process. The agent learns trajectory-aware policies that account for how present solver choices influence downstream error accumulation, while minimizing computational cost under a user-prescribed accuracy tolerance enforced through a Lagrangian reward with online multiplier adaptation. Across sampled 0D homogeneous reactor conditions, the RL-adaptive policy achieves a mean speedup of approximately $3\times$, with speedups ranging from $1.11\times$ to $10.58\times$, while maintaining accurate ignition delays and species profiles for a 106-species \textit{n}-dodecane mechanism and adding approximately $1\%$ inference overhead. Without retraining, the 0D-trained policy transfers to 1D counterflow diffusion flames over strain rates $10$--$2000~\mathrm{s}^{-1}$, delivering consistent $\approx 2.2\times$ speedup relative to CVODE while preserving near-reference temperature accuracy and selecting CVODE at only $12$--$15\%$ of space-time points. Overall, the results demonstrate the potential of the proposed reinforcement learning framework to learn problem-specific integration strategies while respecting accuracy constraints, thereby opening a pathway toward adaptive, self-optimizing workflows for multiphysics systems with spatially heterogeneous stiffness.

强化学习化学模拟数值积分自适应算法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。