arXiv:2507.08977cs.LGcs.AI2025-07被引 6

用机制模拟生成数据训练神经网络,提升科学预测能力与可解释性。

Simulation as Supervision: Mechanistic Pretraining for Scientific Discovery

  • 以多种机制模拟数据预训练神经网络,内化系统动态作为结构先验。
  • 在新冠疫情死亡率预测中,性能接近CDC模型的三倍,且对模型误设鲁棒。
  • 适合需要高可靠性与可解释性的跨学科科学建模任务。

科学建模面临机制理论可解释性与机器学习预测能力之间的权衡。现有混合方法虽将领域知识作为功能约束融入机器学习,但依赖精确数学表达式,当底层方程部分未知或错误时,强约束会引入偏差,阻碍模型从数据中学习。本文提出仿真基神经网络(SGNN),通过机制模拟生成的合成数据对神经网络进行预训练,使模型在多样化的模型结构与真实观测噪声下内化系统动态规律。我们在流行病学、生态学、社会科学和化学等多个领域评估了SGNN。在预测任务中,其性能超越标准数据驱动基线及物理约束混合模型;在新冠死亡率预测中,其预报技能接近美国疾控中心(CDC)平均模型的三倍,并能准确预测高维生态系统。此外,即使在训练数据基于错误假设生成的情况下,SGNN仍表现良好。本框架还引入‘回溯模拟归因’方法,通过识别模拟语料库中最相似的案例来解释真实世界动态。我们证明,多样化机制模拟可作为鲁棒科学推断的有效训练数据。

原文摘要 · Abstract (English)

Scientific modeling faces a tradeoff between the interpretability of mechanistic theory and the predictive power of machine learning. While existing hybrid approaches have made progress by incorporating domain knowledge into machine learning methods as functional constraints, they can be limited by a reliance on precise mathematical specifications. When the underlying equations are partially unknown or misspecified, enforcing rigid constraints can introduce bias and hinder a model's ability to learn from data. We introduce Simulation-Grounded Neural Networks (SGNNs), a framework that incorporates scientific theory by using mechanistic simulations as training data for neural networks. By pretraining on diverse synthetic corpora that span multiple model structures and realistic observational noise, SGNNs internalize the underlying dynamics of a system as a structural prior. We evaluated SGNNs across multiple disciplines, including epidemiology, ecology, social science, and chemistry. In forecasting tasks, SGNNs outperformed both standard data-driven baselines and physics-constrained hybrid models. They nearly tripled the forecasting skill of the average CDC models in COVID-19 mortality forecasts and accurately forecasted high-dimensional ecological systems. SGNNs demonstrated robustness to model misspecification, performing well even when trained on data with incorrect assumptions. Our framework also introduces back-to-simulation attribution, a method for mechanistic interpretability that explains real-world dynamics by identifying their most similar counterparts within the simulated corpus. By unifying these techniques into a single framework, we demonstrate that diverse mechanistic simulations can serve as effective training data for robust scientific inference.

科学建模机制学习仿真训练可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。