arXiv:2604.22168cs.LGcs.SY2026-04

用强化学习优化数字孪生的纠错决策,平衡精度与维护成本。

Optimal sequential decision-making for error propagation mitigation in digital twins

论文配图:Optimal sequential decision-making for error propagation mitigation in digital twins
图 1 · 摘自论文原文
  • 将错误状态建模为马尔可夫决策过程,以最小化维护代价并保持系统精度。
  • 在真实噪声下POMDP策略仍能实现接近MDP 95%的性能。
  • 适合关注数字孪生可靠性与智能运维的研究者和工程师。

本文将模块化数字孪生中的误差传播缓解问题建模为序列决策过程。基于前期研究中通过隐马尔可夫模型(HMM)从代理物理残差中推断潜在错误状态的方法,构建了马尔可夫决策过程(MDP),其中推断出的状态作为状态空间,纠正干预作为动作,奖励函数综合考虑系统保真度与维护成本的权衡。基础转移矩阵由HMM学习参数获得。进一步扩展为部分可观测马尔可夫决策过程(POMDP),通过贝叶斯滤波更新信念分布,并以HMM混淆矩阵作为观测模型,以应对状态分类不完美问题。两种形式均通过动态规划求解,并通过Gillespie随机模拟验证。对比两种无模型强化学习算法(Q-learning与REINFORCE)表明,无需显式模型知识即可学习有效策略。系统性比较显示,MDP策略在累积奖励和正常运行时间占比上表现最优;在现实观测噪声下,POMDP性能约为MDP的95%。对观测质量、修复概率和折扣因子的敏感性分析确认结论稳健,策略层级间差异在 $p < 0.001$ 水平显著。MDP与POMDP之间的性能差距量化了信息的价值,为提升分类精度的投资提供理论依据。

原文摘要 · Abstract (English)

Here, we explore the problem of error propagation mitigation in modular digital twins as a sequential decision process. Building on a companion study that used a Hidden Markov Model (HMM) to infer latent error regimes from surrogate-physics residuals, we develop a Markov Decision Process (MDP) in which the inferred regimes serve as states, corrective interventions serve as actions, and a scalar reward that takes into consideration the cost-benefit tradeoff between system fidelity and maintenance expense. The baseline transition matrix is extracted from the HMM-learned parameters. We then extend the formulation to a Partially Observable MDP (POMDP) that accounts for the imperfect nature of regime classification by maintaining a belief distribution updated via Bayesian filtering, with the HMM confusion matrix serving as the observation model. Both formulations are solved via dynamic programming and validated through Gillespie stochastic simulation. We then benchmark two model-free reinforcement learning algorithms, Q-learning and REINFORCE, to assess whether effective policies can be learned without explicit model knowledge. A systematic comparison of different intervention policies demonstrates that the MDP policy achieves the highest cumulative reward and fraction of time in nominal operation, while the POMDP recovers approximately 95\% of MDP performance under realistic observation noise. Sensitivity analyses across observation quality, repair probability, and discount factor confirm the robustness of these conclusions, and the major gaps in the policy hierarchy are statistically significant at $p < 0.001$. The gap between MDP and POMDP performance quantifies the value of information providing a principled criterion for investing in improved classification accuracy.

数字孪生强化学习决策优化误差传播

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。