用预训练模型提升天文警报分类性能,效率更高且效果更优。
Pre-training vision models for the classification of alerts from wide-field time-domain surveys
- 采用在星系图像上预训练的标准视觉架构
- 性能超越传统定制CNN,推理耗时更少、内存更低
- 适合天文数据科学家快速构建高效分类模型
现代大视场时域巡天通过图像差分生成警报并分发给科研社区,用于研究瞬变、可变及移动天体现象。过去十余年,机器学习工具已广泛应用于此类数据,卷积神经网络(CNN)因其直接从图像预测的特性而被普遍采用。然而,计算机视觉领域近年飞速发展,标准化架构(如ImageNet预训练模型)已成为主流。相比之下,时域天文学仍多使用自定义的CNN并从零训练。本文探索了不同预训练策略与标准架构对警报分类性能的影响。结果表明,采用预训练模型的方案性能达到或超过传统定制化CNN。尤其在银河系动物园(Galaxy Zoo)图像上预训练时表现最佳,优于ImageNet预训练或从零训练。此外,标准化架构设计更优,虽参数更多,但推理时间与内存消耗显著降低。在时空遗产巡天等新巡天项目来临之际,这些发现倡导天文学界采用计算机视觉最新实践,实现更高性能与更高效率。
原文摘要 · Abstract (English)
Modern wide-field time-domain surveys facilitate the study of transient, variable and moving phenomena by conducting image differencing and relaying alerts to their communities. Machine learning tools have been used on data from these surveys and their precursors for more than a decade, and convolutional neural networks (CNNs), which make predictions directly from input images, saw particularly broad adoption through the 2010s. Since then, continually rapid advances in computer vision have transformed the standard practices around using such models. It is now commonplace to use standardized architectures pre-trained on large corpora of everyday images (e.g., ImageNet). In contrast, time-domain astronomy studies still typically design custom CNN architectures and train them from scratch. Here, we explore the effects of adopting various pre-training regimens and standardized model architectures on the performance of alert classification. We find that the resulting models match or outperform a custom, specialized CNN like what is typically used for filtering alerts. Moreover, our results show that pre-training on galaxy images from Galaxy Zoo tends to yield better performance than pre-training on ImageNet or training from scratch. We observe that the design of standardized architectures are much better optimized than the custom CNN baseline, requiring significantly less time and memory for inference despite having more trainable parameters. On the eve of the Legacy Survey of Space and Time and other image-differencing surveys, these findings advocate for a paradigm shift in the creation of vision models for alerts, demonstrating that greater performance and efficiency, in time and in data, can be achieved by adopting the latest practices from the computer vision field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。