arXiv:2505.14877cs.LGastro-ph.SR2025-05被引 1

用自调节机制提升变星分类模型在偏差数据下的可靠性。

A self-regulated convolutional neural network for classifying variable stars

  • 构建生成模型与分类器动态交互,生成物理合理合成光变曲线。
  • 在有偏差的数据上,分类准确率显著优于传统方法。
  • 适合处理标注不均衡、存在数据偏见的天文时序数据场景。

过去二十年,机器学习模型在变星分类中广泛应用,尤其深度学习架构如卷积神经网络、循环神经网络和变换器模型表现优异。然而这些模型需高质量、代表性数据及大量标签样本才能良好泛化,这在时域巡天中常难以实现。数据偏差常导致模型学习并强化训练数据中的固有偏见,而这种问题在仅从同一目录中抽取子集验证时不易察觉。现有研究对变星数据中的偏见问题关注不足,尚无有效解决方案。本文提出一种自调节训练方法,通过物理增强的潜在空间变分自编码器生成合成样本,引入来自盖亚数据发布3的六个物理参数。该方法实现分类器与生成模型的动态交互:生成模型产生针对性合成光变曲线,缓解分类器训练时的混淆,并填补物理参数空间中欠采样区域。多种场景下的实验表明,该自调节训练方法在有偏差数据集上的变星分类性能显著优于传统方法,效果具有统计学意义。

原文摘要 · Abstract (English)

Over the last two decades, machine learning models have been widely applied and have proven effective in classifying variable stars, particularly with the adoption of deep learning architectures such as convolutional neural networks, recurrent neural networks, and transformer models. While these models have achieved high accuracy, they require high-quality, representative data and a large number of labelled samples for each star type to generalise well, which can be challenging in time-domain surveys. This challenge often leads to models learning and reinforcing biases inherent in the training data, an issue that is not easily detectable when validation is performed on subsamples from the same catalogue. The problem of biases in variable star data has been largely overlooked, and a definitive solution has yet to be established. In this paper, we propose a new approach to improve the reliability of classifiers in variable star classification by introducing a self-regulated training process. This process utilises synthetic samples generated by a physics-enhanced latent space variational autoencoder, incorporating six physical parameters from Gaia Data Release 3. Our method features a dynamic interaction between a classifier and a generative model, where the generative model produces ad-hoc synthetic light curves to reduce confusion during classifier training and populate underrepresented regions in the physical parameter space. Experiments conducted under various scenarios demonstrate that our self-regulated training approach outperforms traditional training methods for classifying variable stars on biased datasets, showing statistically significant improvements.

变星分类自调节生成模型天文数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。