用Transformer模型从TESS数据中自动识别引力透镜事件,准确率高达97%。
Microlensify: a Transformer Based Machine Learning Classifier for Microlensing Events Trained on TESS Light Curves

- 基于Transformer的变分自编码器,融合物理规律与真实观测数据训练。
- 在560万条光变曲线中发现0.036%至1.89%为透镜候选,预测持续时间误差极小。
- 可识别伪信号如变星、小行星过境,适合高精度巡天中的透镜探测任务。
引力透镜能揭示难以观测的暗淡致密天体。凌日系外行星巡天卫星(TESS)虽以探测凌日系外行星为主,但具备近全天空覆盖和高采样率的优势。本文利用TESS数据,结合传统方法与机器学习技术,搜索引力透镜候选事件并识别高采样率巡天中的误报。所提出的Microlensify模型是一个物理信息引导的Transformer-based变分自编码器,基于模拟单透镜引力透镜光变曲线与实测TESS第12扇区数据训练。该模型可分类事件、重建光变曲线并估计事件持续时间。应用于约560万条TESS光变曲线后,检测到0.036%至1.89%为透镜候选。通过透镜检测标准及与SIMBAD数据库交叉匹配,最终获得候选列表,并识别出长周期变星、米拉变星、激变变星、红巨星及暂现源等误报。还发现由小行星穿越引起的高斯型峰,是高采样率巡天中潜在的误报来源。模型对事件持续时间的预测准确率可达R² = 0.97。在多个地面巡天发表事件上的测试表明,92.7%被正确识别为引力透镜事件,证明其跨巡天、跨采样率的适用性。
原文摘要 · Abstract (English)
Microlensing can reveal populations of faint compact objects that are otherwise difficult to detect. Depending on their design, all-sky surveys have the potential to search for these objects across the sky. The Transiting Exoplanet Survey Satellite (TESS), primarily designed to detect transiting exoplanets, also provides near all-sky coverage with high cadence. In this work, we use TESS data to search for microlensing candidates using both traditional and machine-learning methods and to identify associated false positives in high-cadence surveys. Microlensify is a physics-informed, transformer-based variational autoencoder trained on simulated single-lens microlensing light curves and real TESS Sector 12 data. The model classifies events, reconstructs light curves, and estimates microlensing event durations. Applied to $\sim 5.6$ million TESS light curves, it identified between $0.036\%$ and $1.89\%$ as microlensing candidates across different TESS pipelines. After applying microlensing detection metrics and cross-matching with SIMBAD, we obtained a final list of candidates and identified false positives including long-period variables, Mira variables, cataclysmic variables, red giants, and transients. We also found Gaussian-like peaks caused by asteroid crossings, a potential source of false positives in high-cadence microlensing surveys. The model also predicts event duration with an accuracy of $R^2 = 0.97$. The model was further tested on published events from different ground-based microlensing surveys, confirming 92.7% as microlensing, demonstrating its applicability across surveys with different cadences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。