用无监督学习与卷积神经网络分类星系氢谱线,提升数据处理效率。
Classification of HI Galaxy Profiles Using Unsupervised Learning and Convolutional Neural Networks: A Comparative Analysis and Methodological Cases of Studies
- 融合聚类与CNN,从318条星系氢谱线中提取特征并分类
- 2D谱线变换模型显著提升分类准确率,优于传统方法
- 方法可推广至SKA等未来射电巡天项目,代码开源
氢是宇宙中最丰富的元素,对理解星系形成与演化至关重要。21 cm中性氢(HI)谱线能映射星系内气体运动,揭示相互作用、结构及恒星形成过程。随着新射电仪器的投入使用,数据量与复杂度持续上升。本文提出一种结合无监督学习与卷积神经网络的框架,用于高效分析和分类积分式HI谱线。研究选取了318条来自CIG星系样本和30,780条来自Arecibo Legacy Fast ALFA Survey(ALFALFA)的数据,通过Busyfit包及多项式、高斯、双洛伦兹模型迭代拟合进行预处理。采用K-means、谱聚类、DBSCAN、层次聚类等方法进行特征提取,并以K-NN、SVM、随机森林分类器优化分类性能,最终通过CNN提升准确率。此外,引入三种基于变换与归一化的2D谱线模型,量化谱线不对称性。该方法在前人关于孤立星系星际介质分析的研究基础上验证有效,有望为正在建设中的平方公里阵列(SKA)提供可复用的数据分析范式。所有代码、模型与数据均已按FAIR原则公开共享。
原文摘要 · Abstract (English)
Hydrogen, the most abundant element in the universe, is crucial for understanding galaxy formation and evolution. The 21 cm neutral atomic hydrogen - HI spectral line maps the gas kinematics within galaxies, providing key insights into interactions, galactic structure, and star formation processes. With new radio instruments, the volume and complexity of data is increasing. To analyze and classify integrated HI spectral profiles in a efficient way, this work presents a framework that integrates Machine Learning techniques, combining unsupervised methods and CNNs. To this end, we apply our framework to a selected subsample of 318 spectral HI profiles of the CIG and 30.780 profiles from the Arecibo Legacy Fast ALFA Survey catalogue. Data pre-processing involved the Busyfit package and iterative fitting with polynomial, Gaussian, and double-Lorentzian models. Clustering methods, including K-means, spectral clustering, DBSCAN, and agglomerative clustering, were used for feature extraction and to bootstrap classification we applied K-NN, SVM, and Random Forest classifiers, optimizing accuracy with CNN. Additionally, we introduced a 2D model of the profiles to enhance classification by adding dimensionality to the data. Three 2D models were generated based on transformations and normalised versions to quantify the level of asymmetry. These methods were tested in a previous analytical classification study conducted by the Analysis of the Interstellar Medium in Isolated Galaxies group. This approach enhances classification accuracy and aims to establish a methodology that could be applied to data analysis in future surveys conducted with the Square Kilometre Array (SKA), currently under construction. All materials, code, and models have been made publicly available in an open-access repository, adhering to FAIR principles.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。