arXiv:2601.05868math.OCcs.LG2026-01被引 1

用强化学习加速无限维参数的实验设计,100倍提速且能发现上游追踪策略

Sequential Bayesian Optimal Experimental Design in Infinite Dimensions via Policy Gradient Reinforcement Learning

  • 将实验设计转为马尔可夫决策过程,用策略梯度学习可复用的设计策略
  • 相比有限元方法提速约100倍,且比随机布点效果更好
  • 适合需要高效多传感器布局的反问题求解场景

针对由偏微分方程(PDE)控制的逆问题,无限维随机场参数下的序贯贝叶斯最优实验设计(SBOED)计算成本高,需在嵌套的贝叶斯反演与设计循环中反复求解前向和伴随PDE。本文将SBOED建模为有限时域马尔可夫决策过程,通过策略梯度强化学习(PGRL)学习可复用的设计策略,实现无需重复求解优化问题的在线设计选择。为提升训练与奖励评估的可扩展性,结合主动子空间投影(参数维度缩减)与主成分分析(状态维度缩减),并引入改进的导数信息感知潜在注意力神经算子(LANO)代理模型,同时预测参数到解映射及其雅可比矩阵。采用基于拉普拉斯近似的D-最优性奖励,框架亦支持如KL散度等期望信息增益目标。进一步提出基于特征值的评估策略,以先验样本作为最大后验(MAP)点的代理,避免重复求解MAP,同时保持准确的信息增益估计。在污染物源追踪的多传感器序贯布设实验中,相较高保真有限元方法实现约100倍加速,性能优于随机布点,且学习到的策略具有物理可解释性,揭示了‘上游’追踪机制。

原文摘要 · Abstract (English)

Sequential Bayesian optimal experimental design (SBOED) for PDE-governed inverse problems is computationally challenging, especially for infinite-dimensional random field parameters. High-fidelity approaches require repeated forward and adjoint PDE solves inside nested Bayesian inversion and design loops. We formulate SBOED as a finite-horizon Markov decision process and learn an amortized design policy via policy-gradient reinforcement learning (PGRL), enabling online design selection from the experiment history without repeatedly solving an SBOED optimization problem. To make policy training and reward evaluation scalable, we combine dual dimension reduction -- active subspace projection for the parameter and principal component analysis for the state -- with an adjusted derivative-informed latent attention neural operator (LANO) surrogate that predicts both the parameter-to-solution map and its Jacobian. We use a Laplace-based D-optimality reward while noting that, in general, other expected-information-gain utilities such as KL divergence can also be used within the same framework. We further introduce an eigenvalue-based evaluation strategy that uses prior samples as proxies for maximum a posteriori (MAP) points, avoiding repeated MAP solves while retaining accurate information-gain estimates. Numerical experiments on sequential multi-sensor placement for contaminant source tracking demonstrate approximately $100\times$ speedup over high-fidelity finite element methods, improved performance over random sensor placements, and physically interpretable policies that discover an ``upstream'' tracking strategy.

实验设计强化学习偏微分方程反问题

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。