用神经隐式表示压缩高能物理稀疏数据,兼顾速度与精度。
Efficient Compression of Sparse Accelerator Data Using Implicit Neural Representations and Importance Sampling
- 用隐式神经网络学习粒子轨迹的连续表示
- 在10^-6至10%稀疏度下实现接近传统算法的压缩率
- 重要性采样加速训练,适合实时高吞吐数据处理
核物理和高能物理中的大型粒子对撞机生成的数据速率极高,可达每秒1太字节至数拍字节。追踪探测器数据的一个显著特点是粒子轨迹在空间中极度稀疏,占用率介于约10^-6至10%之间。此外,下游任务通常更需要连续表示而非基于体素的离散表示,因为信号本身具有连续特性。为此,我们提出一种基于隐式神经表示的数据学习与压缩新方法,并引入重要性采样技术以加速网络训练。该方法在压缩性能上可与传统算法(如MGARD、SZ、ZFP)相媲美,同时通过重要性采样策略实现显著提速,且保持可忽略的精度损失。
原文摘要 · Abstract (English)
High-energy, large-scale particle colliders in nuclear and high-energy physics generate data at extraordinary rates, reaching up to $1$ terabyte and several petabytes per second, respectively. The development of real-time, high-throughput data compression algorithms capable of reducing this data to manageable sizes for permanent storage is of paramount importance. A unique characteristic of the tracking detector data is the extreme sparsity of particle trajectories in space, with an occupancy rate ranging from approximately $10^{-6}$ to $10\%$. Furthermore, for downstream tasks, a continuous representation of this data is often more useful than a voxel-based, discrete representation due to the inherently continuous nature of the signals involved. To address these challenges, we propose a novel approach using implicit neural representations for data learning and compression. We also introduce an importance sampling technique to accelerate the network training process. Our method is competitive with traditional compression algorithms, such as MGARD, SZ, and ZFP, while offering significant speed-ups and maintaining negligible accuracy loss through our importance sampling strategy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。