轻量级模型NEXUS可跨域迁移,提升粒子物理等科学任务精度。
A Lightweight Foundation Model for Collider Physics with Multi-Domain Adaptation

- 用300万参数的全连接自编码器,无监督预训练于大型对撞机数据。
- 在小标注数据下,对撞机任务的精度优于从零训练的同类模型。
- 模型计算轻量,适合实时或边缘部署,适用于引力波、洪水预测等场景。
我们提出一种轻量级基础建模方法(NEXUS),利用对撞机物理数据的预训练能力,实现向其他科学数据集的跨域任务迁移。该模型采用约300万参数的全连接自编码器,在大型对撞机碰撞数据集上进行无监督预训练,基于电荷粒子轨迹特征建模。下游对撞机分析任务(如运动学回归与事件分类)在使用预训练权重时,仅需少量标注数据即可达到更高精度,显著优于从头训练的同类架构。通过潜在空间解析,进一步验证了预训练收益,并拓展至引力波、洪水预测与神经活动等其他领域。相比同规模的Transformer方法,NEXUS展现出更优的计算效率,为科学实验中的高效推理与实时/边缘应用提供了可能。
原文摘要 · Abstract (English)
We present a lightweight approach to foundation modeling (\textbf{NEXUS}) that leverages pre-trained learning from collider physics data towards out-of-domain tasks in other scientific datasets, using a fully connected autoencoder model with approximately 3 million parameters. The model pre-trains with no supervision over a large-scale collision dataset from the Large Hadron Collider modeled by charged particle track features. Downstream tasks for collider analyses, such as kinematic regression and event classification, are developed on pre-trained model weights and achieve improved accuracy with only small labeled datasets when compared to equivalent architectures trained from scratch. The benefits of pre-training are additionally investigated through latent space interpretation and application to other domains, including gravitational waves, flood forecasting, and neural activity. Furthermore, the relative computational simplicity of NEXUS is demonstrated compared to transformer approaches at comparable scale, opening the door to power-efficient inference and real-time or edge applications of foundation models in scientific experiments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。