用主动学习提升稀有行星宜居性分类效率,减少标注需求。
Active Learning for Planet Habitability Classification under Extreme Class Imbalance
- 基于不确定性的主动学习策略,动态选择最有价值的待标注行星。
- 仅需少量标注样本即可逼近监督模型性能,标签效率显著提升。
- 适合天文领域数据稀缺、标签不平衡的场景,支持保守优先级排序。
随着系外行星目录规模扩大和异质性增加,系统评估宜居性面临挑战,尤其因潜在宜居行星极度稀少且标签持续演化。本文构建了来自宜居世界目录与NASA系外行星档案的统一数据集,将宜居性评估建模为二分类问题。采用梯度提升决策树建立监督基线,并优化召回率以优先识别稀有宜居行星。该模型嵌入池采样主动学习框架,比较基于不确定性的边界采样与随机查询在多轮实验和不同标注预算下的表现。结果表明,主动学习显著降低达到监督模型性能所需的标注实例数,实现明显标签效率提升。为对接实际天文学应用,我们对独立训练的主动学习模型进行集成,利用平均概率与不确定性对原非宜居行星进行排序,识别出一个稳健候选体,展示主动学习如何支持保守、带不确定性的后续目标优先排序,而非推测性再分类。结果表明,主动学习为标签不平衡、信息不全、观测资源有限的数据场景提供了合理引导框架。
原文摘要 · Abstract (English)
The increasing size and heterogeneity of exoplanet catalogs have made systematic habitability assessment challenging, particularly given the extreme scarcity of potentially habitable planets and the evolving nature of their labels. In this study, we explore the use of pool-based active learning to improve the efficiency of habitability classification under realistic observational constraints. We construct a unified dataset from the Habitable World Catalog and the NASA Exoplanet Archive and formulate habitability assessment as a binary classification problem. A supervised baseline based on gradient-boosted decision trees is established and optimized for recall in order to prioritize the identification of rare potentially habitable planets. This model is then embedded within an active learning framework, where uncertainty-based margin sampling is compared against random querying across multiple runs and labeling budgets. We find that active learning substantially reduces the number of labeled instances required to approach supervised performance, demonstrating clear gains in label efficiency. To connect these results to a practical astronomical use case, we aggregate predictions from independently trained active-learning models into an ensemble and use the resulting mean probabilities and uncertainties to rank planets originally labeled as non-habitable. This procedure identifies a single robust candidate for further study, illustrating how active learning can support conservative, uncertainty-aware prioritization of follow-up targets rather than speculative reclassification. Our results indicate that active learning provides a principled framework for guiding habitability studies in data regimes characterized by label imbalance, incomplete information, and limited observational resources.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。