针对粒子探测数据稀疏性,提出可变速率神经压缩方法
Variable Rate Neural Compression for Sparse Detector Data
- 用稀疏卷积识别关键点,动态调整压缩策略
- 重建精度提升75%,压缩比提高10%,模型小200倍
- 适合高稀疏度的实时探测数据压缩场景
高能大型对撞机产生极高速率的数据流。为降低数据量并满足存储带宽需求,开发实时高吞吐数据压缩算法至关重要。在新建成的相对论重离子对撞机sPHENIX实验中,时间投影室(TPC)作为主要追踪探测器,记录气体腔体内三维粒子轨迹。质子-质子碰撞下数据占用率可低至$10^{-3}$,这种稀疏性使传统无学习损失压缩算法(如SZ、ZFP、MGARD)面临挑战。相比之下,基于深度学习的模型,特别是使用卷积神经网络的压缩模型,在压缩比和重建精度上已超越传统方法。然而,现有研究仍缺乏对这类模型处理稀疏数据集(如对撞机数据)有效性的系统评估。此外,多数深度学习模型无法根据数据稀疏性自适应调节处理速度,影响效率。为此,我们提出一种新方法:通过稀疏卷积实现关键点识别,用于TPC数据压缩。所提算法BCAE-VS相比前代最优模型,重建精度提升75%,压缩比提高10%,且模型规模缩小超过两个数量级。实验验证表明,随着数据稀疏性增加,模型吞吐量也随之提升。
原文摘要 · Abstract (English)
High-energy large-scale particle colliders generate data at extraordinary rates. Developing real-time high-throughput data compression algorithms to reduce data volume and meet the bandwidth requirement for storage has become increasingly critical. Deep learning is a promising technology that can address this challenging topic. At the newly constructed sPHENIX experiment at the Relativistic Heavy Ion Collider, a Time Projection Chamber (TPC) serves as the main tracking detector, which records three-dimensional particle trajectories in a volume of a gas-filled cylinder. In terms of occupancy, the resulting data flow can be very sparse reaching $10^{-3}$ for proton-proton collisions. Such sparsity presents a challenge to conventional learning-free lossy compression algorithms, such as SZ, ZFP, and MGARD. In contrast, emerging deep learning-based models, particularly those utilizing convolutional neural networks for compression, have outperformed these conventional methods in terms of compression ratios and reconstruction accuracy. However, research on the efficacy of these deep learning models in handling sparse datasets, like those produced in particle colliders, remains limited. Furthermore, most deep learning models do not adapt their processing speeds to data sparsity, which affects efficiency. To address this issue, we propose a novel approach for TPC data compression via key-point identification facilitated by sparse convolution. Our proposed algorithm, BCAE-VS, achieves a $75\%$ improvement in reconstruction accuracy with a $10\%$ increase in compression ratio over the previous state-of-the-art model. Additionally, BCAE-VS manages to achieve these results with a model size over two orders of magnitude smaller. Lastly, we have experimentally verified that as sparsity increases, so does the model's throughput.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。