arXiv:2606.19329astro-ph.IMcs.LG2026-06中稿 · The Astrophysical …

用机器学习解决天文图像中X射线与光学源的匹配难题

The Chandra-Gaia Catalog of Counterparts: Resolving ambiguous Gaia matches to X-ray sources in the Chandra Source Catalog using Machine Learning

论文配图:The Chandra-Gaia Catalog of Counterparts: Resolving ambiguous Gaia matches to X-ray sources in the Chandra Source Catalog using Machine Learning
图 1 · 摘自论文原文
  • 结合光度、颜色和距离信息,用梯度提升树模型识别真实对应体
  • 为11.3万条X射线源找到光学对应体,发现7000组多重候选
  • 适用于多源交叉匹配,尤其适合高密度区域的精准匹配

本文提出一种框架,将钱德拉源目录(CSC v2.1)中的X射线源与盖亚数据发布3(Gaia DR3)中的光学源进行交叉匹配。不同于传统仅依赖位置的方法,该方法利用星等、颜色和距离等属性识别真实对应体,检测偶然重合,并解决多个候选者共存时的歧义问题。通过基于贝叶斯的NWAY框架构建高置信度训练集,训练了轻量级梯度提升分类器(LightGBM)。在约25.4万条独立X射线源中,成功匹配到约11.3万条对应体,其中约7000条存在多个合理候选。另有约2万条在传统距离匹配下有匹配结果,但本方法未找到对应体,其中一半归因于偶然重合。在钱德拉猎户座超深场项目(COUP)上验证,该方法在不使用位置信息的情况下复现了95%的NWAY匹配结果。我们发布了包含约11.3万条匹配结果、约7000条备选匹配及约2万条模糊关联的目录,支持未来双波段源的统计研究。文中讨论了局限性,并提出了可推广至其他交叉匹配场景的通用框架。

原文摘要 · Abstract (English)

We present a framework to cross-match sources from the Chandra Source Catalog (CSC v2.1) with optical sources from Gaia Data Release 3. Unlike purely spatial approaches, we use source properties such as magnitudes, colors, and distances to identify true counterparts, detect chance coincidences, and resolve ambiguities when multiple plausible candidates exist. We define a training set of high-confidence matches using NWAY, a Bayesian cross-matching framework that accounts for positional errors and source densities. We train a gradient-boosted classifier (LightGBM) on a variety of features from both catalogs. Of the ~$254$k unique X-ray sources, we find counterparts for ~$113$k sources, of which plausible multiple counterparts are found for ~$7$k. We find no counterparts for ~$20$k sources for which separation-based cross-matching does find a match, and attribute half of these to chance coincidences. We validate the pipeline on the Chandra Orion Ultradeep Project (COUP), where the machine-learning matches reproduce 95% of NWAY cross-matches without using any positional information. We release a catalog of the ~$113$k Chandra-Gaia counterparts, together with ~$7$k alternative matches and ~$20$k ambiguous NWAY associations, supporting future population studies of sources detectable by both Chandra and Gaia. We discuss limitations and provide a generalization of the framework that is applicable in other cross-matching scenarios.

天文交叉匹配机器学习天体物理数据融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。