arXiv:2603.28804physics.ins-detcs.LG2026-03被引 2

用专家混合与高效微调构建可扩展的量能器模拟基础模型

Generalizable Foundation Models for Calorimetry via Mixtures-of-Experts and Parameter Efficient Fine Tuning

  • 基于下一令牌预测框架,采用专家混合预训练+轻量微调
  • 支持新材料、新粒子类型增量添加,避免灾难性遗忘
  • 适合高能物理实验中持续迭代的探测器模拟需求

现代粒子物理实验面临日益增长的高保真探测器模拟需求,随着亮度提升,计算资源已逼近极限。深度生成模型正作为传统蒙特卡洛模拟的替代方案,受大语言模型和下一词预测范式启发。本文提出一种基于下一词变换器架构的量能器通用基础模型,支持材料、粒子种类和探测器配置的模块化适配。通过混合专家预训练与参数高效微调策略,实现无灾难性遗忘的可控、增量式模型扩展:预训练骨干网络在多种吸收材料上生成电磁簇射,新材料通过新增并微调轻量级专家模块实现;新粒子类型则借助参数高效微调与模块化词表完成扩展,保持基础模型完整性。该设计支持随新数据集到来持续集成知识,契合真实探测器开发流程。此外,在标准大语言模型优化下,该模型计算效率优于传统生成方法。结果表明,下一词架构为可扩展、物理感知的量能器基础模型提供了可行路径,适用于未来高能物理实验。

原文摘要 · Abstract (English)

Modern particle physics experiments face an increasing demand for high-fidelity detector simulation as luminosities rise and computational requirements approach the limits of available resources. Deep generative models have emerged as promising surrogates for traditional Monte Carlo simulation, with recent advances drawing inspiration from large language models (LLM) and next-token prediction paradigms. In this work, we introduce a generalizable foundation model for calorimetry built on next-token transformer backbones, designed to support modular adaptation across materials, particle species, and detector configurations. Our approach combines Mixture-of-Experts pre-training with parameter-efficient fine-tuning strategies to enable controlled, additive model expansion without catastrophic forgetting. A pre-trained backbone is trained to generate electromagnetic showers across multiple absorber materials, while new materials are incorporated through the addition and tuning of lightweight expert modules. Extensions to new particle types are achieved via parameter-efficient fine-tuning and modular vocabularies, preserving the integrity of the base model. This design enables efficient, incremental knowledge integration as new simulation datasets become available, a critical requirement in realistic detector-development workflows. In addition, we demonstrate that next-token calorimeter models are computationally competitive with standard generative approaches under established LLM optimization procedures. These results establish next-token architectures as a viable path toward extensible, physics-aware foundation models for calorimetry and future high-energy physics experiments.

量能器模拟基础模型专家混合高效微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。