arXiv:2510.19168astro-ph.COastro-ph.IM2025-10中稿 · NeurIPS

用标准模型预训练,减少新物理模拟需求,但需防负迁移。

Transfer Learning Beyond the Standard Model

  • 在ΛCDM上预训练,再微调到新物理场景,降低仿真成本。
  • 仅需少量新物理仿真即可实现准确推断,提升效率。
  • 瓶颈结构效果最佳,但强参数混淆会导致学习失败。

机器学习可实现强大的宇宙学推断,但通常需要大量覆盖多种宇宙学模型的高保真模拟。迁移学习通过跨模型复用知识,有望降低模拟成本。我们发现,在标准宇宙学模型ΛCDM上进行预训练,并在各种超越ΛCDM的场景(包括大质量中微子、修改引力、原初非高斯性)上微调,可显著减少对超越ΛCDM模拟的需求。然而,当ΛCDM与超越ΛCDM参数之间存在强烈物理退化时,也可能出现负迁移。我们测试了多种迁移架构,发现引入瓶颈结构表现最优。研究揭示了物理领域基础模型方法的机遇与风险:预训练可加速推断,但也可能阻碍新物理的学习。

原文摘要 · Abstract (English)

Machine learning enables powerful cosmological inference but typically requires many high-fidelity simulations covering many cosmological models. Transfer learning offers a way to reduce the simulation cost by reusing knowledge across models. We show that pre-training on the standard model of cosmology, $Λ$CDM, and fine-tuning on various beyond-$Λ$CDM scenarios -- including massive neutrinos, modified gravity, and primordial non-Gaussianities -- can enable inference with significantly fewer beyond-$Λ$CDM simulations. However, we also show that negative transfer can occur when strong physical degeneracies exist between $Λ$CDM and beyond-$Λ$CDM parameters. We consider various transfer architectures, finding that including bottleneck structures provides the best performance. Our findings illustrate the opportunities and pitfalls of foundation-model approaches in physics: pre-training can accelerate inference, but may also hinder learning new physics.

迁移学习宇宙学预训练物理建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。