融合仿真与实测数据,提升代理模型精度与可靠性
Bayesian Surrogate Training on Multiple Data Sources: A Hybrid Modeling Strategy
- 分源训练后融合预测分布,或单模型联合优化
- 在合成与真实案例中均提升预测准确率与覆盖率
- 适用于需诊断模拟模型缺陷的工程建模场景
代理模型常作为复杂仿真模型的计算高效近似,支持反问题求解、敏感性分析和概率前向预测等任务。然而,仿真模型本身是对现实系统的简化,可能遗漏关键过程或存在输入/边界条件误设。这些线索可能隐藏在真实测量数据中,但传统方法忽略此类信息。本文提出两种新颖的概率化方法,在代理模型训练中融合仿真数据与实测数据:第一种为各数据源分别训练代理模型并组合其预测分布;第二种则通过单一代理模型联合训练。两种方法均采用独立于代理模型族的新型加权策略整合异构数据源。通过合成与真实案例研究,展示了两种方法在概念上的差异与优势。结果表明,该方法能有效提升预测精度、覆盖范围,并诊断底层仿真模型的问题,有助于深化系统认知与未来模型改进。
原文摘要 · Abstract (English)
Surrogate models are often used as computationally efficient approximations to complex simulation models, enabling tasks such as solving inverse problems, sensitivity analysis, and probabilistic forward predictions, which would otherwise be computationally infeasible. During training, surrogate parameters are fitted such that the surrogate reproduces the simulation model's outputs as closely as possible. However, the simulation model itself is merely a simplification of the real-world system, often missing relevant processes or suffering from misspecifications e.g., in inputs or boundary conditions. Hints about these might be captured in real-world measurement data, and yet, we typically ignore those hints during surrogate building. In this paper, we propose two novel probabilistic approaches to integrate simulation data and real-world measurement data during surrogate training. The first method trains separate surrogate models for each data source and combines their predictive distributions, while the second incorporates both data sources by training a single surrogate. Both hybrid modeling approaches employ a novel weighting strategy for combining heterogeneous data sources during surrogate training, which operates independently of the chosen surrogate family. We show the conceptual differences and benefits of the two approaches through both synthetic and real-world case studies. The results demonstrate the potential of these methods to improve predictive accuracy, predictive coverage, and to diagnose problems in the underlying simulation model. These insights can improve system understanding and future model development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。