arXiv:2606.25923cs.LG2026-06

让数字孪生更懂决策:优化政策排序而非单纯模拟误差

$\text{DT}^2$: Decision-Targeted Digital Twins

  • 用拟合Q值评估策略,生成排序监督信号
  • 训练时保留策略间相对优劣关系,提升决策准确性
  • 适合需要可靠政策选择的工业场景应用

数字孪生(DT)是现实系统的虚拟模型,可通过模拟不同策略下的情景辅助决策。然而,传统基于机器学习的数字孪生并未针对此目标优化。我们证明,在模型容量受限时,仅最小化单步转移误差会导致政策排序性能下降。实验也表明,即使使用表达力强的模型,这一问题依然存在。为此,我们提出决策导向的数字孪生训练范式 $ ext{DT}^2$:首先利用拟合Q值评估从离线数据中获取候选策略的价值,再训练数字孪生生成能保持这些策略对之间相对排序的仿真轨迹,采用与架构无关的损失函数。我们在多种设置和架构下验证了方法的有效性,$ ext{DT}^2$ 在策略排序准确性和策略选择中的决策遗憾方面均优于传统训练方式,且对训练内与未见策略均有提升,同时维持良好的原始仿真保真度。

原文摘要 · Abstract (English)

A digital twin (DT) is a virtual model of a real-world system that can assist decision-making by simulating scenarios induced by different policies. However, typical machine learning-based DTs do not optimise for this use case. We prove that, when model capacity is limited, training DTs to minimise one-step transition errors can produce suboptimal models for ranking sets of policies according to a reward function. We further show that this holds empirically, even with expressive model classes. To address this, we introduce $\text{DT}^2$, a decision-targeted DT training paradigm. Firstly, $\text{DT}^2$ uses fitted Q-evaluation to estimate values of candidate policies from offline data. A DT is then trained to generate rollouts that preserve pairwise policy rankings derived from these proxy ground-truth values with an architecture-agnostic loss function. We empirically demonstrate the efficacy of our method across a range of settings and architectures. $\text{DT}^2$ consistently improves policy ranking and reduces decision regret during policy selection relative to conventional DT training, both for policies used during training and for unseen policies, while maintaining a good level of raw simulation fidelity.

数字孪生决策优化强化学习策略排序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。