首个天文变星光变曲线基准,评估时序大模型在不规则数据上的表现
StarEmbed: Benchmarking Time Series Foundation Models on Astronomical Observations of Variable Stars
- 构建包含4万颗恒星的光变曲线公开基准,覆盖7类专家标注数据
- Chronos模型虽未在天文数据上预训练,但在聚类和异常检测上达顶尖水平
- 适用于研究时序大模型泛化能力或天体物理分类任务的研究者
当前时序基础模型(TSFM)训练数据多忽略不规则采样等复杂情况。天文领域中恒星光度时间序列数据量巨大,具有不规则采样、多变量和异方差性特征。我们提出StarEmbed,首个基于真实观测的光变曲线公开基准,包含约40,000颗恒星的专家标注数据,涵盖七类,并支持聚类、分类及分布外(OOD)源检测评估。我们对不同架构与训练策略的TSFMs及领域专用Transformer进行了评测。结果表明,尽管Chronos系列模型仅在非天文、规则采样数据上预训练,其在光变曲线聚类和OOD检测上仍达到当前最优(SOTA)性能;虽然尚无TSFM在分类任务上全面超越传统领域基线,但展现出优异的泛化能力。StarEmbed推动了通用光变曲线嵌入的发展,并为提升复杂数据下的时序模型性能提供新路径。
原文摘要 · Abstract (English)
Current time series foundation model (TSFM) training corpora largely omit data with certain complexities like irregular temporal sampling. Astronomical time series of stellar fluxes (light curves) are available in immense quantities and exhibit irregular sampling, multiple variates, and heteroskedasticity. We introduce StarEmbed, the first public benchmark for light curves comprised of real observations of ~40,000 stars expert-labeled across seven classes and evaluations in clustering, classification, and out-of-distribution (OOD) source detection. We benchmark TSFMs with differing architecture and training strategies as well as domain-specific transformers. Our results demonstrate that the Chronos family, despite being pre-trained on regularly sampled non-astronomical data, yields state-of-the-art (SOTA) performance in light curve clustering and OOD detection. While no TSFM strictly surpasses the classification performance of the long-established domain baseline, they do demonstrate excellent generalization abilities. StarEmbed marks a step toward universal light curve embeddings and improved TSFM performance on challenging data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。