用多模态自监督学习提升天文学图像与光变、光谱数据的联合分析能力
AstroM$^3$: A self-supervised multimodal model for astronomy
- 扩展CLIP模型至光变、光谱与元数据三模态,实现跨模态自监督预训练
- 在少量标注数据下分类准确率提升12.6%,光变数据分类达91.5%
- 可发现新天体类型,支持异常检测与相似性搜索等下游任务
尽管机器学习模型已广泛用于天文研究,但模型输入通常局限于单一数据源(如图像或时间序列)或部分元数据。随着宽视场、多路复用观测资源的普及,目标天体常具备多种观测模式。本文构建了天文多模态数据集,并提出AstroM³,一种自监督预训练方法,使模型能同时学习时间序列测光、光谱和天体物理元数据。具体地,将CLIP模型扩展至三模态设置。在微调的监督任务中,结果表明CLIP预训练显著提升时间序列测光分类性能,准确率从84.6%提高至91.5%;在标签数据有限时,准确率最高提升12.6%,证明利用大规模无标签数据的有效性。除分类外,所学嵌入还可用于误分类识别、相似性搜索与异常检测。令人意外的是,通过流形学习与降维算法,模型“重新发现”了造父变星亚型及两类旋转变星亚型。据我们所知,这是天文学中首个n>2模态模型,该方法可自然拓展至n>3模态。
原文摘要 · Abstract (English)
While machine-learned models are now routinely employed to facilitate astronomical inquiry, model inputs tend to be limited to a primary data source (namely images or time series) and, in the more advanced approaches, some metadata. Yet with the growing use of wide-field, multiplexed observational resources, individual sources of interest often have a broad range of observational modes available. Here we construct an astronomical multimodal dataset and propose AstroM$^3$, a self-supervised pre-training approach that enables a model to learn from multiple modalities simultaneously. Specifically, we extend the CLIP (Contrastive Language-Image Pretraining) model to a trimodal setting, allowing the integration of time-series photometry data, spectra, and astrophysical metadata. In a fine-tuning supervised setting, our results demonstrate that CLIP pre-training improves classification performance for time-series photometry, where accuracy increases from 84.6% to 91.5%. Furthermore, CLIP boosts classification accuracy by up to 12.6% when the availability of labeled data is limited, showing the effectiveness of leveraging larger corpora of unlabeled data. In addition to fine-tuned classification, we can use the trained model in other downstream tasks that are not explicitly contemplated during the construction of the self-supervised model. In particular we show the efficacy of using the learned embeddings for misclassifications identification, similarity search, and anomaly detection. One surprising highlight is the "rediscovery" of Mira subtypes and two Rotational variable subclasses using manifold learning and dimension reduction algorithm. To our knowledge this is the first construction of an $n>2$ mode model in astronomy. Extensions to $n>3$ modes is naturally anticipated with this approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。