构建天文数据压缩基准集,推动神经网络提升科学观测数据传输效率
AstroCompress: A benchmark dataset for multi-purpose compression of astronomical data
- 设计四个新数据集加一个旧数据集,覆盖太空/地面、多波段、时序等模式
- 神经压缩方法在16位图像上实现比传统算法更高压缩率,提升数据采集能力
- 为科研机构提供可复现的压缩评测工具,适合天文与数据压缩研究者
空间与地面天文台因寒冷黑暗的理想观测条件而选址偏远,导致数据传输能力受限。这一瓶颈直接影响可观测数据量,在现代望远镜成本高昂的背景下,任何无损压缩技术的改进都可能带来数十亿美元级的额外科学产出。传统无损压缩依赖人工设计,而神经数据压缩可通过端到端学习,利用天文图像独特的时空与波长结构,超越经典方法。本文提出AstroCompress:一个面向天体物理数据的神经压缩挑战基准,包含四个新数据集和一个遗留数据集,均采用16位无符号整型图像,涵盖空间、地面、多波段及时间序列成像模式。我们提供代码以便捷访问数据,并基准测试七种无损压缩方法(三种神经、四种非神经,包括所有实用的前沿算法)。实验表明,无损神经压缩能有效提升观测站的数据收集能力,并为科学应用中的神经压缩采纳提供指导。尽管本研究聚焦无损压缩,未来亦可探索有损压缩的可能性。
原文摘要 · Abstract (English)
The site conditions that make astronomical observatories in space and on the ground so desirable -- cold and dark -- demand a physical remoteness that leads to limited data transmission capabilities. Such transmission limitations directly bottleneck the amount of data acquired and in an era of costly modern observatories, any improvements in lossless data compression has the potential scale to billions of dollars worth of additional science that can be accomplished on the same instrument. Traditional lossless methods for compressing astrophysical data are manually designed. Neural data compression, on the other hand, holds the promise of learning compression algorithms end-to-end from data and outperforming classical techniques by leveraging the unique spatial, temporal, and wavelength structures of astronomical images. This paper introduces AstroCompress: a neural compression challenge for astrophysics data, featuring four new datasets (and one legacy dataset) with 16-bit unsigned integer imaging data in various modes: space-based, ground-based, multi-wavelength, and time-series imaging. We provide code to easily access the data and benchmark seven lossless compression methods (three neural and four non-neural, including all practical state-of-the-art algorithms). Our results on lossless compression indicate that lossless neural compression techniques can enhance data collection at observatories, and provide guidance on the adoption of neural compression in scientific applications. Though the scope of this paper is restricted to lossless compression, we also comment on the potential exploration of lossy compression methods in future studies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。