用自监督+微调的Transformer自动编码器,从重离子碰撞数据中提取高效特征。
Latent Representation Learning in Heavy-Ion Collisions with MaskPoint Transformer
- 两阶段训练:先自监督预训练,再监督微调,直接从无标签数据学潜空间表示
- 在区分大/小碰撞系统任务上,分类准确率显著高于PointNet
- 能捕捉可观测量之外的非线性关联,适合研究夸克-胶子等离子体性质
高能核物理中的核心挑战是如何从重离子碰撞(HIC)的高维末态数据中提取有效特征,以支持可靠的下游分析。传统方法依赖选定的可观测量,可能遗漏数据中微妙但物理相关的结构。为此,我们提出一种基于Transformer的自编码器,采用两阶段范式:自监督预训练后接监督微调。预训练编码器直接从无标签的HIC数据中学习潜表示,构建紧凑且信息丰富的特征空间,可适配多种物理任务。作为案例研究,该方法用于区分大与小碰撞系统,分类准确率显著优于PointNet。主成分分析与SHAP解释进一步表明,该自编码器捕捉了超越单个可观测量的复杂非线性相关性,生成具有强区分力和可解释性的特征。这些结果证明该两阶段框架是HIC特征学习的一般且稳健基础,为研究夸克-胶子等离子体性质及其他涌现现象开辟新路径。代码已公开于https://github.com/Giovanni-Sforza/MaskPoint-AMPT。
原文摘要 · Abstract (English)
A central challenge in high-energy nuclear physics is to extract informative features from the high-dimensional final-state data of heavy-ion collisions (HIC) in order to enable reliable downstream analyses. Traditional approaches often rely on selected observables, which may miss subtle but physically relevant structures in the data. To address this, we introduce a Transformer-based autoencoder trained with a two-stage paradigm: self-supervised pre-training followed by supervised fine-tuning. The pretrained encoder learns latent representations directly from unlabeled HIC data, providing a compact and information-rich feature space that can be adapted to diverse physics tasks. As a case study, we apply the method to distinguish between large and small collision systems, where it achieves significantly higher classification accuracy than PointNet. Principal component analysis and SHAP interpretation further demonstrate that the autoencoder captures complex nonlinear correlations beyond individual observables, yielding features with strong discriminative and explanatory power. These results establish our two-stage framework as a general and robust foundation for feature learning in HIC, opening the door to more powerful analyses of quark--gluon plasma properties and other emergent phenomena. The implementation is publicly available at https://github.com/Giovanni-Sforza/MaskPoint-AMPT.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。