arXiv:2410.00271astro-ph.COastro-ph.IM2024-10被引 8

构建首个大规模星系多模态数据集,助力机器学习在天文观测中的应用。

GalaxiesML: a dataset of galaxy images, photometry, redshifts, and structural parameters for machine learning

  • 整合28万余张星系图像与光度、红移等参数,统一标注真实红移。
  • 使用图像比仅用光度估算红移的误差降低10倍,精度显著提升。
  • 适合天文机器学习研究者,尤其关注下一代巡天数据建模者。

本文介绍一个专为机器学习设计的星系数据集,包含286,401个星系的图像、光度、光谱红移及结构参数,源自Hyper-Suprime-Cam Survey PDR2的五波段(g, r, i, z, y)观测,所有红移均有光谱确认。该数据集均匀一致,异常值极少,信噪比覆盖真实范围,适合训练机器学习模型。数据集红移范围为0.01至4,峰值在1.5,超过2.5后迅速衰减。我们展示了利用图像进行红移估计的案例:在红移0.1至1.25区间,图像方法的红移估计偏差比仅用光度低一个数量级。此数据集将推动欧几里得(Euclid)和大型巡天望远镜(LSST)等下一代巡天的数据分析方法发展。

原文摘要 · Abstract (English)

We present a dataset built for machine learning applications consisting of galaxy photometry, images, spectroscopic redshifts, and structural properties. This dataset comprises 286,401 galaxy images and photometry from the Hyper-Suprime-Cam Survey PDR2 in five imaging filters ($g,r,i,z,y$) with spectroscopically confirmed redshifts as ground truth. Such a dataset is important for machine learning applications because it is uniform, consistent, and has minimal outliers but still contains a realistic range of signal-to-noise ratios. We make this dataset public to help spur development of machine learning methods for the next generation of surveys such as Euclid and LSST. The aim of GalaxiesML is to provide a robust dataset that can be used not only for astrophysics but also for machine learning, where image properties cannot be validated by the human eye and are instead governed by physical laws. We describe the challenges associated with putting together a dataset from publicly available archives, including outlier rejection, duplication, establishing ground truths, and sample selection. This is one of the largest public machine learning-ready training sets of its kind with redshifts ranging from 0.01 to 4. The redshift distribution of this sample peaks at redshift of 1.5 and falls off rapidly beyond redshift 2.5. We also include an example application of this dataset for redshift estimation, demonstrating that using images for redshift estimation produces more accurate results compared to using photometry alone. For example, the bias in redshift estimate is a factor of 10 lower when using images between redshift of 0.1 to 1.25 compared to photometry alone. Results from dataset such as this will help inform us on how to best make use of data from the next generation of galaxy surveys.

星系数据机器学习红移估计天文数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。