arXiv:2608.21756cs.LGcs.CV2026-08

研究发现神经网络架构比训练更能决定探测器数据表征的可复用性。

How Architecture and Training Affect TPC Representations Across Experiments

论文配图:How Architecture and Training Affect TPC Representations Across Experiments
图 1 · 摘自论文原文
  • 用冻结编码器加探针分析架构与训练对表征的影响
  • 随机初始化的PointNet结构在多个任务中仍具高信息量
  • 不同架构组织嵌入空间方式不同,但跨探测器通用性良好

深度学习正趋向基础模型范式,在实验物理中,这使得模型与学习到的表征可在原开发实验之外复用。本文通过在冻结编码器上使用探针,评估表征在不同实验与探测器系统间的可复用性。这些探针在下游微调前揭示了任务相关结构,弥补了仅靠微调的不足。结合随机权重对照组,可分离出架构与编码器训练的贡献。时间投影室(TPC)数据作为测试平台,因其事件可表示为变长稀疏张量,且探测器几何、事件拓扑和科学任务差异显著。我们研究固定维度的TPC事件表征是否可跨分类任务、实验和探测器系统复用。使用Sparse ResNet与PointNet风格编码器,从GADGET II TPC与AT-TPC的四个数据集生成512维嵌入。随机初始化编码器用于隔离训练前架构的贡献。随后在分类任务上训练编码器,冻结参数后为每个下游任务训练线性或非线性探针。结果表明,架构带来的结构在跨实验与探测器系统间依然有效。随机初始化的PointNet式表征在多个任务中信息量显著。两种架构以不同方式组织嵌入空间,但均未表现出大规模系统性性能下降。结果说明,架构是TPC嵌入中任务相关结构的主要来源,应被明确纳入表征学习评估与可复用探测器模型开发中。

原文摘要 · Abstract (English)

Deep-learning efforts have increasingly shifted toward foundation model approaches. In experimental physics, this allows models and learned representations to be reused beyond the experiments in which they were developed. This work evaluates the reusability of representations across experiments and detector systems using probes on frozen encoders. These probes reveal task-relevant structure before downstream adaptation, complementing fine-tuning. Together with random-weight controls, they distinguish contributions from architecture and encoder training that downstream performance alone cannot resolve. Time projection chamber (TPC) data provide a useful testbed because events from TPC systems can be represented as variable-length sparse tensors, while detector geometries, event topologies, and scientific tasks can differ substantially. We investigate whether fixed-dimensional TPC event representations can be reused across classification tasks, experiments, and detector systems. Sparse ResNet and PointNet-style encoders produce 512-dimensional embeddings for four datasets from the GADGET II TPC and AT-TPC. Randomly initialized encoders isolate the contribution from architecture before supervised training. We then train each encoder on a classification task, freeze its parameters, and train a linear or nonlinear probe for each downstream task. We find that this architecture-induced structure remains useful across experiments and detector systems. The randomly initialized PointNet-style representation is highly informative on several tasks. The two architectures organize their embedding spaces differently, but neither exhibits a large, systematic loss of utility cross-detector. These results show that architecture is a major source of task-relevant structure in TPC embeddings and should be treated explicitly when assessing representation learning and developing reusable detector models.

表征学习探测器建模架构影响

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。