arXiv:2509.17645astro-ph.EPastro-ph.IM2025-09

RAVEN用机器学习自动判断系外行星候选是否为假阳性,准确率超91%。

RAVEN: RAnking and Validation of ExoplaNets

  • 结合贝叶斯框架与梯度提升树、高斯过程分类器进行多场景假阳性评估。
  • 在1361个独立样本上实现91%准确率,假阳性识别AUC超97%。
  • 开源云应用,适合天文团队快速验证TESS系外行星候选体。

我们提出RAVEN,一个针对TESS系外行星候选体的新一代筛选与验证流程。该流程采用贝叶斯框架,通过梯度提升决策树与高斯过程分类器,基于包含8种天体假阳性情景的合成数据集,计算候选体为行星的后验概率。训练集涵盖模拟行星及真实系统噪声和恒星活动导致的非模拟假阳性。结合位置先验等信息,最终后验概率超过99%且半径小于8$R_{igoplus}$的候选体被认定为统计验证通过。本版本适用于经TESS科学处理中心发布的光变曲线,周期0.5至16天、掩食深度大于300ppm的候选体。所有假阳性场景的AUC均高于97%,仅一种低于99%。在1361个预分类的TOI独立样本中,整体准确率达91%;当概率阈值设为0.9时,精确率为97%,召回率为66%。RAVEN以云端应用形式公开发布,便于社区使用。

原文摘要 · Abstract (English)

We present RAVEN, a newly developed vetting and validation pipeline for TESS exoplanet candidates. The pipeline employs a Bayesian framework to derive the posterior probability of a candidate being a planet against a set of False Positive (FP) scenarios, through the use of a Gradient Boosted Decision Tree and a Gaussian Process classifier, trained on comprehensive synthetic training sets of simulated planets and 8 astrophysical FP scenarios injected into TESS lightcurves. These training sets allow large scale candidate vetting and performance verification against individual FP scenarios. A Non-Simulated FP training set consisting of real TESS candidates caused primarily by stellar variability and systematic noise is also included. The machine learning derived probabilities are combined with scenario specific prior probabilities, including the candidates' positional probabilities, to compute the final posterior probabilities. Candidates with a planetary posterior probability greater than 99% against each FP scenario and whose implied planetary radius is less than 8$R_{\oplus}$ are considered to be statistically validated by the pipeline. In this first version, the pipeline has been developed for candidates with a lightcurve released from the TESS Science Processing Operations Centre, an orbital period between 0.5 and 16 days and a transit depth greater than 300ppm. The pipeline obtained area-under-curve (AUC) scores > 97% on all FP scenarios and > 99% on all but one. Testing on an independent external sample of 1361 pre-classified TOIs, the pipeline achieved an overall accuracy of 91%, demonstrating its effectiveness for automated ranking of TESS candidates. For a probability threshold of 0.9 the pipeline reached a precision of 97% with a recall score of 66% on these TOIs. The RAVEN pipeline is publicly released as a cloud-hosted app, making it easily accessible to the community.

系外行星机器学习数据验证TESS

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。