用强化学习从噪声观测中直接学未知动态的模型,无需真实状态数据。
Reinforcement learning based data assimilation for unknown state model
- 将参数估计转为马尔可夫决策过程,用强化学习求解最优模型参数。
- 在高维场景下比传统方法更准确、更抗噪,无需真实状态训练数据。
- 适合动态方程未知、观测含噪声的复杂系统状态估计任务。
数据同化(DA)在众多应用中已成为状态估计的关键工具,但当底层动力学方程未知时极具挑战性。现有机器学习方法通常依赖预生成的无噪声训练数据构建代理状态转移模型,而实际中难以获取无噪声的真实状态序列。为此,本文提出一种新方法,将强化学习与基于集合的贝叶斯滤波相结合,直接从噪声观测中学习未知动力学的代理状态转移模型,无需使用真实状态轨迹。具体地,将代理模型参数的最大似然估计过程建模为离散时间马尔可夫决策过程(MDP),学习代理模型等价于寻找该MDP的最优策略,可通过强化学习有效求解。模型离线训练完成后,可在在线阶段使用基于学习动力学的滤波方法进行状态估计。该框架适用于非线性及部分可观测测量模型等多种观测场景。数值实验表明,所提方法在高维设置下具有更高的准确性和鲁棒性。
原文摘要 · Abstract (English)
Data assimilation (DA) has increasingly emerged as a critical tool for state estimation across a wide range of applications. It is significantly challenging when the governing equations of the underlying dynamics are unknown. To this end, various machine learning approaches have been employed to construct a surrogate state transition model in a supervised learning framework, which relies on pre-computed training datasets. However, it is often infeasible to obtain noise-free ground-truth state sequences in practice. To address this challenge, we propose a novel method that integrates reinforcement learning with ensemble-based Bayesian filtering methods, enabling the learning of surrogate state transition model for unknown dynamics directly from noisy observations, without using true state trajectories. Specifically, we treat the process for computing maximum likelihood estimation of surrogate model parameters as a sequential decision-making problem, which can be formulated as a discrete-time Markov decision process (MDP). Under this formulation, learning the surrogate transition model is equivalent to finding an optimal policy of the MDP, which can be effectively addressed using reinforcement learning techniques. Once the model is trained offline, state estimation can be performed in the online stage using filtering methods based on the learned dynamics. The proposed framework accommodates a wide range of observation scenarios, including nonlinear and partially observed measurement models. A few numerical examples demonstrate that the proposed method achieves superior accuracy and robustness in high-dimensional settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。