arXiv:2510.12958astro-ph.IMastro-ph.HE2025-10

用模拟数据预训练,少样本也能精准分析天文时间序列。

Simulation-Based Pretraining and Domain Adaptation for Astronomical Time Series with Minimal Labeled Data

  • 用多个天文巡天的模拟数据预训练模型,学习通用特征表示。
  • 仅需少量真实数据微调,分类、红移估计等任务性能显著提升。
  • 跨领域泛化能力强,零样本迁移对新望远镜数据有效。

天文时间序列分析面临标注数据稀缺的瓶颈。本文提出一种基于模拟数据的预训练方法,大幅减少对真实观测标注数据的需求。模型在多个天文巡天(ZTF 和 LSST)的模拟数据上训练,学习可迁移的通用表征,微调时仅需极少真实数据即可在分类、红移估计和异常检测任务中超越基线方法。令人惊讶的是,仅在现有望远镜(ZTF)数据上训练的模型,即可实现对未来的望远镜(LSST)模拟数据的零样本迁移,性能相当。此外,尽管训练数据为暂现事件,模型仍能有效泛化至完全不同现象(如开普勒望远镜的变星),展现强大跨域能力。该方法为标注数据匮乏但可构建模拟场景的领域提供了实用建模方案。

原文摘要 · Abstract (English)

Astronomical time-series analysis faces a critical limitation: the scarcity of labeled observational data. We present a pre-training approach that leverages simulations, significantly reducing the need for labeled examples from real observations. Our models, trained on simulated data from multiple astronomical surveys (ZTF and LSST), learn generalizable representations that transfer effectively to downstream tasks. Using classifier-based architectures enhanced with contrastive and adversarial objectives, we create domain-agnostic models that demonstrate substantial performance improvements over baseline methods in classification, redshift estimation, and anomaly detection when fine-tuned with minimal real data. Remarkably, our models exhibit effective zero-shot transfer capabilities, achieving comparable performance on future telescope (LSST) simulations when trained solely on existing telescope (ZTF) data. Furthermore, they generalize to very different astronomical phenomena (namely variable stars from NASA's \textit{Kepler} telescope) despite being trained on transient events, demonstrating cross-domain capabilities. Our approach provides a practical solution for building general models when labeled data is scarce, but domain knowledge can be encoded in simulations.

时间序列模拟预训练零样本迁移天文学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。