用视觉语言模型识别中微子事件,提升分类准确率与可解释性。
Adapting Vision-Language Models for Neutrino Event Classification in High-Energy Physics
- 用微调的LLaMA 3.2构建视觉语言模型处理像素化探测器数据。
- 相比CNN和ViT,VLM在分类准确率和鲁棒性上均表现更优。
- 适合需要可解释性推理的高能物理实验场景。
大型语言模型在处理结构化与非结构化数据方面展现出强大能力。本文探索将视觉语言模型(VLM)应用于高能物理实验中的中微子相互作用识别任务,采用微调版LLaMA 3.2模型处理像素化探测器数据。我们将其与先进卷积神经网络(CNN)及同架构的视觉变换器(ViT-h/14)进行对比,评估其分类性能与预测可解释性。结果表明,基于Transformer的架构在分类准确率和鲁棒性上优于传统CNN;而融合文本/语义信息的VLM进一步提升了灵活性,支持基于推理的可解释预测。这表明大规模Transformer模型,特别是视觉语言模型,具备作为物理事件分类通用骨干的潜力,兼具高性能、强鲁棒性与可解释性,为实验中微子物理中的多模态推理开辟新路径。
原文摘要 · Abstract (English)
Recent advances in Large Language Models (LLMs) have demonstrated their remarkable capacity to process and reason over structured and unstructured data modalities beyond natural language. In this work, we explore the applications of Vision Language Models (VLMs), specifically a fine-tuned variant of LLaMA 3.2 to the task of identifying neutrino interactions in pixelated detector data from high-energy physics (HEP) experiments. We benchmark this model against a state-of-the-art convolutional neural network (CNN) architecture, similar to those used in major neutrino experiments, which have achieved high efficiency and purity in classifying electron and muon neutrino events, and also a Vision Transformer (ViT-h/14), which is the same architecture inside the VLM's vision encoder. Our evaluation considers both classification performance and interpretability of the model predictions, comparing a VLM with a vision-only transformer (ViT) and a convolutional neural network (CNN) baseline. We find that transformer-based architectures outperform conventional CNNs in classification accuracy and robustness, with the VLM providing additional flexibility through the integration of auxiliary textual or semantic information and enabling more interpretable, reasoning-based predictions. These results highlight the potential of large transformer models, particularly vision-language models, as general-purpose backbones for physics event classification, combining strong performance, robustness, and interpretability, and opening new avenues for multimodal reasoning in experimental neutrino physics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。