PRISM可动态控制复杂度,高效从海量模型中选最优解
Scalable Simulation-Based Model Inference with Test-Time Complexity Control
- 用编码器-解码器架构联合推断模型结构与参数
- 测试时调节先验实现复杂度可控,支持超大规模模型筛选
- 适用于生物物理建模等需多模型比较的科学场景
仿真在科学发现中至关重要。许多应用的瓶颈已不再是运行仿真器,而是从大量可能的仿真器中选择——每个对应一个与观测一致的前向模型/假设。在大规模模型族中,经典贝叶斯模型选择方法不实用。此外,现有近似模型选择方法通常在训练时硬编码固定先验或复杂度惩罚,要求用户在看到数据前就确定简约性假设。我们提出PRISM,一种基于仿真的编码器-解码器框架,可联合推断离散模型结构与连续参数的后验分布,并通过在测试时调节网络所依赖的可调模型先验,实现模型复杂度的动态控制。我们在合成符号回归任务上展示了PRISM可扩展至含数十亿个模型实例的组合型模型家族。作为科学应用,我们在扩散MRI数据的生物物理建模中评估了PRISM,证明其可在合成和真实脑成像数据上,对多个多组分模型进行有效模型选择。
原文摘要 · Abstract (English)
Simulation plays a central role in scientific discovery. In many applications, the bottleneck is no longer running a simulator; it is choosing among large families of plausible simulators, each corresponding to different forward models/hypotheses consistent with observations. Over large model families, classical Bayesian workflows for model selection are impractical. Furthermore, amortized model selection methods typically hard-code a fixed model prior or complexity penalty at training time, requiring users to commit to a particular parsimony assumption before seeing the data. We introduce PRISM, a simulation-based encoder-decoder that infers a joint posterior over both discrete model structures and associated continuous parameters, while enabling test-time control of model complexity via a tunable model prior that the network is conditioned on. We show that PRISM scales to families with combinatorially many (up to billions) of model instantiations on a synthetic symbolic regression task. As a scientific application, we evaluate PRISM on biophysical modeling for diffusion MRI data, showing the ability to perform model selection across several multi-compartment models, on both synthetic and in vivo neuroimaging data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。