arXiv:2606.16219cs.CEcs.LG2026-06

从观测数据中自动发现关键变量,构建可解释的简化随机模型。

Graphical conditional generative modeling for digital twin modeling

论文配图:Graphical conditional generative modeling for digital twin modeling
图 1 · 摘自论文原文
  • 通过条件生成模型识别影响目标变量全条件分布的输入变量
  • 在多种系统中实现与完整模型相当的性能,且模型更简洁可解释
  • 适合需要可解释性与鲁棒性的工业建模、控制与强化学习场景

数字孪生建模在模型不确定性下进行控制与数据融合时,常面临开放式的保真度问题:增加变量、数据流和时间尺度会无限提升模型复杂度,导致难以维护、验证、解释,也无法用于压力或安全测试。为此,我们提出一种方法,从观测数据中发现仅需描述相关量的简约随机代理模型。该方法通过识别哪些候选输入影响目标量的完整条件分布(而非仅条件均值)来筛选变量。这一区分对随机、粗粒度或部分观测系统至关重要,因依赖关系可能体现在方差、尾部行为、多模态或不确定性变化上,而非确定性函数关系。框架结合条件生成建模(学习目标变量的条件分布)与基于高斯过程的方差分析(通过核函数分解),实现非关键输入的迭代剔除与可解释结构发现。在控制场景中,所得代理模型可视为学习到的马尔可夫决策过程:不仅识别转移模型,还确定使动态有效马尔可夫所需的态变量、动作变量与记忆变量。在涉及随机动力系统、缺失变量、偏微分方程控制、强化学习及经济数据的多个案例中,所发现结构生成的可解释随机代理模型,在下游任务中的表现与使用全部变量训练的模型相当。

原文摘要 · Abstract (English)

Digital twin modeling, including control and data assimilation under model uncertainty, often faces an open-ended fidelity problem: adding variables, data streams, and time scales can indefinitely increase model complexity, ultimately producing systems that are difficult to maintain, validate, interpret, and use for stress or safety testing. As an alternative, one can seek parsimonious stochastic surrogate models built only on the variables needed to describe the relevant quantities of interest. We introduce a framework for discovering such variables from observational data by identifying which candidate inputs influence the full conditional law of a target quantity, rather than only its conditional mean. This distinction is essential in stochastic, coarse-grained, or partially observed systems, where dependencies may appear through changes in variability, tail behavior, multimodality, or uncertainty rather than through deterministic functional relationships. The framework couples conditional generative modeling, which learns the conditional distribution of the target given candidate inputs, with Gaussian-process-based analysis of variance (through kernel mode decomposition), which enables iterative pruning of non-influential inputs and interpretable structure discovery. In control settings, the resulting surrogate can be interpreted as a learned Markov decision process: the method identifies not only a transition model, but also the state, action, and memory variables needed to make the learned dynamics effectively Markovian. Across examples involving stochastic dynamical systems, missing variables, PDE control, reinforcement learning, and economic data, the discovered structures yield interpretable stochastic surrogates whose downstream performance is comparable to models trained on the full variable set.

数字孪生生成建模变量筛选可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。