用合成图像预训练模型,提升星团质量估算精度
Self-supervised Synthetic Pretraining for Inference of Stellar Mass Embedded in Dense Gas
- 用一 million 张合成分形图自监督预训练视觉变压器
- 在有限高分辨率模拟数据上,预测性能优于监督模型
- 可无监督识别星形成区域结构,适合天体物理研究
恒星质量是决定恒星性质与演化的核心参数。然而,在致密气体遮蔽且分布不均的恒星形成区中,传统球对称动力学方法难以准确估计质量。监督学习虽可建模复杂结构与质量关系,但需大量高质量标注数据,而高分辨率磁流体模拟成本高昂。本文采用自监督框架 DINOv2,利用一百万张合成分形图像对视觉变压器进行预训练,随后将冻结的模型应用于有限的高分辨率磁流体模拟数据。结果表明,合成预训练显著提升了冻结特征在恒星质量回归任务中的表现,其性能略优于在相同有限模拟数据上训练的监督模型。主成分分析显示提取特征具有语义意义,表明该模型可在无需标签或微调的情况下实现无监督的星形成区域分割。
原文摘要 · Abstract (English)
Stellar mass is a fundamental quantity that determines the properties and evolution of stars. However, estimating stellar masses in star-forming regions is challenging because young stars are obscured by dense gas and the regions are highly inhomogeneous, making spherical dynamical estimates unreliable. Supervised machine learning could link such complex structures to stellar mass, but it requires large, high-quality labeled datasets from high-resolution magneto-hydrodynamical (MHD) simulations, which are computationally expensive. We address this by pretraining a vision transformer on one million synthetic fractal images using the self-supervised framework DINOv2, and then applying the frozen model to limited high-resolution MHD simulations. Our results demonstrate that synthetic pretraining improves frozen-feature regression stellar mass predictions, with the pretrained model performing slightly better than a supervised model trained on the same limited simulations. Principal component analysis of the extracted features further reveals semantically meaningful structures, suggesting that the model enables unsupervised segmentation of star-forming regions without the need for labeled data or fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。