用稀疏自编码器挖掘中微子模型的可解释隐变量,提升方向预测精度。
Finding and using interpretable latents in a neutrino foundation model with sparse autoencoders

- 通过稀疏自编码器识别模型内部物理概念的可解释隐变量。
- 基于隐变量构建误差预测头,将角度分辨率从20.2°提升至3.2°。
- 适用于希望理解并优化物理模型内部表示的研究者。
我们首次将基于稀疏自编码器的机械可解释性方法应用于粒子物理领域。研究了一个在IceCube数据上预训练并微调用于方向重建的中微子基础模型,通过严格验证流程(包括留出测试、匹配干扰控制及独立字典训练的复现),识别出模型表示中的物理概念图谱。因果干预显示,方向预测头几乎未使用该图谱信息。受此启发,我们在同一事件级表示上训练了一个不确定性头,以预测模型的角度重建误差。与方向头不同,该头因果依赖于图谱中的质量与亮度特征。在20%选择效率下,该可解释估计器将中位角度分辨率从20.2°提升至3.2°。结果表明,机械可解释性可揭示模型内部编码的物理规律,并指导下游任务的设计以有效利用这些信息。
原文摘要 · Abstract (English)
We present a first application of sparse-autoencoder-based mechanistic interpretability to particle physics. Studying a neutrino foundation model pretrained on IceCube data and fine-tuned for direction reconstruction, we identify a validated atlas of physical concepts in the model representation, using a strict validation protocol consisting of held-out tests, matched nuisance controls, and replication across independent dictionary trainings. Causal interventions show that the direction head barely draws on this atlas. Motivated by this underused information, we train an uncertainty head on the same event-level representation to predict the model's angular reconstruction error. Unlike the direction head, it depends causally on quality and brightness features from the atlas. At $20\%$ selection efficiency, this interpretable estimator improves the median angular resolution from $20.2^\circ$ to $3.2^\circ$. These results suggest that mechanistic interpretability can reveal learned latent physics encoded within a model's internal representation and help design downstream tasks that exploit it.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。