arXiv:2509.18990cs.LGmath.DS2025-09被引 4

用模拟数据训练神经网络,让模型在真实数据少时仍能精准预测。

Learning From Simulators: A Theory of Simulation-Grounded Learning

  • 将模拟数据训练的模型视为基于模拟先验的贝叶斯预测器。
  • 在模拟与现实有偏差时,性能下降可被理论刻画,且模型仍能恢复隐藏参数。
  • 提出反向模拟归因法,可解释预测依据,适合科学建模与机制分析。

模拟基神经网络(SGNN)是仅用机制模拟生成的合成数据训练的预测模型,在真实标签稀缺或不可得的领域表现优异,但缺乏理论基础。本文将其置于统一统计框架下:在标准损失函数下,SGNN 可被解释为在模拟诱导先验下的摊销贝叶斯预测器。经验风险最小化可收敛至合成分布下的贝叶斯最优预测。利用经典分布偏移理论,我们刻画了当模拟偏离真实时性能退化的规律。此外,本文还建立了两类特定结果:(i) 在何种条件下未观测科学参数可通过模拟学习;(ii) 提出一种回溯模拟归因方法,通过关联预测相似的模拟实例提供机制解释,并保证后验一致性。数值实验验证了理论预测:SGNN 能恢复潜变量,在模拟失配下保持鲁棒性,且在模型选择任务中误差仅为 AIC 的一半。这些结果确立了 SGNN 作为数据受限环境下科学预测的原理性与实用性框架。

原文摘要 · Abstract (English)

Simulation-Grounded Neural Networks (SGNNs) are predictive models trained entirely on synthetic data from mechanistic simulations. They have achieved state-of-the-art performance in domains where real-world labels are limited or unobserved, but lack a formal underpinning. We place SGNNs in a unified statistical framework. Under standard loss functions, they can be interpreted as amortized Bayesian predictors trained under a simulator-induced prior. Empirical risk minimization then yields convergence to the Bayes-optimal predictor under the synthetic distribution. We employ classical results on distribution shift to characterize how performance degrades when the simulator diverges from reality. Beyond these consequences, we develop SGNN-specific results: (i) conditions under which unobserved scientific parameters are learnable via simulation, and (ii) a back-to-simulation attribution method that provides mechanistic explanations of predictions by linking them to the simulations the model deems similar, with guarantees of posterior consistency. We provide numerical experiments to validate theoretical predictions. SGNNs recover latent parameters, remain robust under mismatch, and outperform classical tools: in a model selection task, SGNNs achieve half the error of AIC in distinguishing mechanistic dynamics. These results establish SGNNs as a principled and practical framework for scientific prediction in data-limited regimes.

模拟训练贝叶斯推理科学建模模型解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。