arXiv:2606.13978astro-ph.IMcs.LG2026-06

用主成分分析压缩光谱数据,实现天文谱型高效分类

Classification of Astronomical Spectra Using PCA-Compressed Flux and Inverse-Variance Features

  • 将光谱通量与不确定性联合建模,通过PCA降维提取特征
  • 使用LightGBM分类器达到94.6%准确率和92.1%平衡准确率
  • 适合天体物理数据处理与自动化分类任务的研究者参考

本文评估了一种用于将斯隆数字巡天数据发布17版(SDSS DR17)天文光谱分类为恒星、星系和类星体的信号处理与监督学习流程。每条光谱由其测量通量和逆方差信息表示,结合光谱形状与波长依赖的可靠性轮廓。在统一的对数波长网格上重采样后,通量与逆方差向量分别进行标准化并使用主成分分析(PCA)压缩。所得主成分拼接后用于训练多个分类器。表现最佳的是LightGBM梯度提升分类器,在测试集上达到94.6%的准确率和92.1%的平衡准确率。

原文摘要 · Abstract (English)

This paper evaluates a signal-processing and supervised-learning pipeline for classifying SDSS DR17 astronomical spectra into stars, galaxies, and quasars. Each spectrum is represented by its measured flux and inverse-variance information, combining spectral shape with a wavelength-dependent reliability profile. After resampling onto a common logarithmic wavelength grid, the flux and inverse-variance vectors are standardized and separately compressed using principal component analysis. The resulting components are concatenated and used to train several classifiers. The best performance was obtained with the LightGBM gradient-boosting classifier, reaching $94.6\%$ accuracy and $92.1\%$ balanced accuracy on the test set.

光谱分类主成分分析天文数据机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。