arXiv:2604.17910cs.AIcs.LG2026-04

用因果结构压缩状态空间,提升工程仿真中约束修复的效率与成功率。

Physics-Informed Causal MDPs for Sequential Constraint Repair in Engineering Simulation Pipelines

论文配图:Physics-Informed Causal MDPs for Sequential Constraint Repair in Engineering Simulation Pipelines
图 1 · 摘自论文原文
  • 基于生命周期排序假设,识别跨层因果关系并压缩高维状态空间。
  • 仅300次训练即达76.2%修复成功率,优于最强基线5.4个百分点。
  • 融合物理先验的估计算法降低方差,适合复杂工程系统优化场景。

在具有大二值状态空间的约束马尔可夫决策过程(CMDP)中,离策略学习面临根本矛盾:因果动态识别需结构假设,而高效策略学习则依赖状态压缩。本文提出PI-CMDP框架,其约束依赖关系在生命周期排序假设(LOA)下构成分层有向无环图(DAG)。提出“识别-压缩-估计”三阶段流程:(i) 识别:在LOA成立时,可通过后门调整识别跨层边权重;当LOA失效时,提供形式化的部分识别边界;(ii) 压缩:在层优先规则与可交换性假设下,将状态基数从2^(WL)降至(W+1)^L;(iii) 估计:引入物理引导的双重稳健估计器,在物理先验优于学习模型时保持无偏并降低方差常数。在TPS基准测试(4,206个轨迹)上,PI-CMDP以仅300次训练达到76.2%修复成功率,较最强基线提升5.4个百分点;全数据条件下仍保持2.8个百分点优势(83.4% vs. 80.6%),且显著降低级联失败率。所有改进在5组独立种子下均显著(配对t检验,p < 0.02)。

原文摘要 · Abstract (English)

Off-policy learning in constrained MDPs with large binary state spaces faces a fundamental tension: causal identification of transition dynamics requires structural assumptions, while sample-efficient policy learning requires state-space compression. We introduce PI-CMDP, a framework for CMDPs whose constraint dependencies form a layered DAG under a Lifecycle Ordering Assumption (LOA). We propose an Identify-Compress-Estimate pipeline: (i) Identify: LOA enables backdoor identification of causal edge weights for cross-layer pairs, with formal partial-identification bounds when LOA is violated; (ii) Compress: a Markov abstraction compresses state cardinality from 2^(WL) to (W+1)^L under layer-priority regularity and exchangeability; and (iii) Estimate: a physics-guided doubly-robust estimator remains unbiased and reduces the variance constant when the physics prior outperforms a learned model. We instantiate PI-CMDP on constraint repair in engineering simulation pipelines. On the TPS benchmark (4,206 episodes), PI-CMDP achieves 76.2% repair success rate with only 300 training episodes versus 70.8% for the strongest baseline (+5.4 pp), narrowing to +2.8 pp (83.4% vs. 80.6%) in the full-data regime, while substantially reducing cascade failure rates. All improvements are consistent across 5 independent seeds (paired t-test p < 0.02).

强化学习因果推理工程优化状态压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。