用机器学习自动识别巡天数据中的星际物体,提升发现效率。
Machine Learning Methods for Automated Interstellar Object Classification with LSST
- 用梯度提升机和随机森林算法区分星际物体与太阳系天体
- 最佳模型准确率、召回率、F1值均超99.8%
- 适合天文数据处理与行星系统研究者参考
薇拉·鲁宾天文台即将开展的时空遗产巡天(LSST)将产生前所未有的太阳系天体数据,包括稀有的星际物体(ISOs)。由于ISOs稀少且观测窗口短暂,加之LSST将生成海量数据,其识别与分类面临巨大挑战。本研究探索了多种机器学习算法在模拟LSST数据中对ISO轨道段(tracklets)的自动化分类性能。实验表明,梯度提升机(GBM)和随机森林(RF)优于随机梯度下降(SGD)和神经网络(NN)。RF分析显示,许多衍生特征(Digest2值)比原始观测值更具判别力。其中,GBM模型表现最优,精度、召回率与F1分数分别达0.9987、0.9986和0.9987。研究成果为构建高效可靠的自动化ISO发现系统奠定基础,有助于深入理解其他行星系统的物质组成与形成过程。该方法有望集成至LSST数据处理流程,提升对这类稀有天体的及时发现与后续观测能力。
原文摘要 · Abstract (English)
The Legacy Survey of Space and Time, to be conducted with the Vera C. Rubin Observatory, is poised to revolutionize our understanding of the Solar System by providing an unprecedented wealth of data on various objects, including the elusive interstellar objects (ISOs). Detecting and classifying ISOs is crucial for studying the composition and diversity of materials from other planetary systems. However, the rarity and brief observation windows of ISOs, coupled with the vast quantities of data to be generated by LSST, create significant challenges for their identification and classification. This study aims to address these challenges by exploring the application of machine learning algorithms to the automated classification of ISO tracklets in simulated LSST data. We employed various machine learning algorithms, including random forests (RFs), stochastic gradient descent (SGD), gradient boosting machines (GBMs), and neural networks (NNs), to classify ISO tracklets in simulated LSST data. We demonstrate that GBM and RF algorithms outperform SGD and NN algorithms in accurately distinguishing ISOs from other Solar System objects. RF analysis shows that many derived Digest2 values are more important than direct observables in classifying ISOs from the LSST tracklets. The GBM model achieves the highest precision, recall, and F1 score, with values of 0.9987, 0.9986, and 0.9987, respectively. These findings lay the foundation for the development of an efficient and robust automated system for ISO discovery using LSST data, paving the way for a deeper understanding of the materials and processes that shape planetary systems beyond our own. The integration of our proposed machine learning approach into the LSST data processing pipeline will optimize the survey's potential for identifying these rare and valuable objects, enabling timely follow-up observations and further characterization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。