arXiv:2508.19376cs.LGcs.AI2025-08

用视觉语言模型分析中微子事件,性能优于传统方法。

Fine-Tuning Vision-Language Models for Neutrino Event Analysis in High-Energy Physics Experiments

  • 基于LLaMA 3.2的视觉语言模型,结合图像与文本上下文进行分类。
  • 在NOvA和DUNE实验数据上,准确率、召回率等指标优于或持平CNN基准。
  • 适合需要多模态信息融合的高能物理事件分析任务。

大型语言模型在跨模态推理方面展现出巨大潜力。本文探索基于LLaMA 3.2的视觉语言模型(VLM)在高能物理实验中对像素化探测器图像中的中微子相互作用进行分类的应用。我们在NOvA和DUNE等实验的基准数据上,对比了该VLM与传统卷积神经网络(CNN)在分类准确率、精确率、召回率及AUC-ROC等指标上的表现。结果表明,VLM不仅在性能上达到或超越现有CNN方法,还能实现更丰富的推理过程,并更好地融合辅助文本或语义信息。这些发现表明,VLM可作为高能物理事件分类的通用骨干模型,为实验中微子物理的多模态研究开辟新路径。

原文摘要 · Abstract (English)

Recent progress in large language models (LLMs) has shown strong potential for multimodal reasoning beyond natural language. In this work, we explore the use of a fine-tuned Vision-Language Model (VLM), based on LLaMA 3.2, for classifying neutrino interactions from pixelated detector images in high-energy physics (HEP) experiments. We benchmark its performance against an established CNN baseline used in experiments like NOvA and DUNE, evaluating metrics such as classification accuracy, precision, recall, and AUC-ROC. Our results show that the VLM not only matches or exceeds CNN performance but also enables richer reasoning and better integration of auxiliary textual or semantic context. These findings suggest that VLMs offer a promising general-purpose backbone for event classification in HEP, paving the way for multimodal approaches in experimental neutrino physics.

中微子视觉语言模型高能物理多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。