用自然语言生成模型并联合估计参数,实现从描述中自动发现合适模型。
Program Synthesis for Simulation-Based Inference: Joint Model Selection and Parameter Estimation
- 通过大语言模型生成候选模拟程序,结合反馈迭代优化。
- 在多类系统上准确识别出符合数据的模型家族,表现随数据信息量变化。
- 适合需要从零构建模型结构的研究者,尤其适用于复杂系统建模。
神经仿真推断可实现对复杂模型的参数估计,但通常要求用户预先指定固定结构的模拟器。本文提出一种联合模型选择与参数估计的框架,结合大语言模型进行程序生成与神经仿真推断。给定系统的自然语言描述和观测数据,LLM生成候选模拟程序,经反馈驱动的变异迭代优化,并通过神经密度估计评估。该方法可在一组候选模型中进行推断,而不仅限于固定模型内的参数。在涵盖确定性动力学、随机流行病模型以及引力透镜图像中暗物质子结构推断的基准测试中,方法能从开放式提示中识别出合理的模型族,其准确性反映数据的信息含量及模型可辨识性。
原文摘要 · Abstract (English)
Neural simulation-based inference enables parameter estimation for complex models, but typically requires the user to specify a simulator encoding a fixed model structure. We present a framework for joint model selection and parameter estimation that combines large language models for program synthesis with neural simulation-based inference. Given a natural language description of the system and data under investigation, an LLM proposes candidate simulator programs which are iteratively refined via feedback-driven mutation and evaluated using neural density estimation. The approach enables simulation-based inference over a pool of models, not just parameters within a fixed model. On benchmarks spanning deterministic dynamics, stochastic epidemic models, and dark matter substructure inference from gravitational-lensing images, the method identifies plausible model families from open-ended prompts, with accuracy that reflects the information content of the data and identifiability of candidate models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。