arXiv:2410.06922astro-ph.EPastro-ph.IM2024-10中稿 · publication in the…

用机器学习从不完整数据中估算系外行星质量,提升预测精度并生成合成行星样本。

Estimating Exoplanet Mass using Machine Learning on Incomplete Datasets

  • 采用kNN×KDE等五种算法处理多维缺失数据,实现行星质量的分布式估计。
  • 数据越多即使不完整,预测性能越优,可对任意已知属性组合的行星进行质量推断。
  • 输出概率分布反映置信度,揭示系外行星群体特征,适合天文数据挖掘研究者。

系外行星档案库是研究系外行星性质的重要资源,但统计分析受限于大量缺失值。其中最具信息量的物理参数之一是行星质量,超过70%的已发现行星缺乏实测质量值。本文比较了五种机器学习算法在处理多维不完整数据集时估算缺失属性的能力,分别基于部分完整六项属性子集与包含六项和八项属性的不完整全集进行测试。结果表明,即便新增数据不完整,增加数据量仍能提升插补效果,并实现对任意已知属性组合的行星进行质量预测。最优算法为新提出的kNN×KDE,可返回被插补属性的概率分布,其形状反映算法置信度,同时揭示系外行星群体的潜在分布特征。通过多个实例展示该方法对凌星法与径向速度法发现的行星的适用性。最后,利用kNN×KDE生成大规模合成行星样本,识别出多维空间中具有特定属性组合的潜在行星类别。所有代码开源。

原文摘要 · Abstract (English)

The exoplanet archive is an incredible resource of information on the properties of discovered extrasolar planets, but statistical analysis has been limited by the number of missing values. One of the most informative bulk properties is planet mass, which is particularly challenging to measure with more than 70\% of discovered planets with no measured value. We compare the capabilities of five different machine learning algorithms that can utilize multidimensional incomplete datasets to estimate missing properties for imputing planet mass. The results are compared when using a partial subset of the archive with a complete set of six planet properties, and where all planet discoveries are leveraged in an incomplete set of six and eight planet properties. We find that imputation results improve with more data even when the additional data is incomplete, and allows a mass prediction for any planet regardless of which properties are known. Our favored algorithm is the newly developed $k$NN$\times$KDE, which can return a probability distribution for the imputed properties. The shape of this distribution can indicate the algorithm's level of confidence, and also inform on the underlying demographics of the exoplanet population. We demonstrate how the distributions can be interpreted with a series of examples for planets where the discovery was made with either the transit method, or radial velocity method. Finally, we test the generative capability of the $k$NN$\times$KDE to create a large synthetic population of planets based on the archive, and identify potential categories of planets from groups of properties in the multidimensional space. All codes are Open Source.

系外行星机器学习数据插补生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。