用跨模态自编码器学习星系图像与光谱的关联,提升数据匮乏下的宇宙研究能力。
Multi-Modal Masked Autoencoders for Learning Image-Spectrum Associations for Galaxy Evolution and Cosmology
- 基于变换器架构,对75%图像和光谱数据进行掩码重建,学习共享表征
- 能恢复星系形状、发射线峰等关键物理特征,红移预测性能优于已有模型
- 适合处理图像多而光谱少的天文大数据场景,为构建天体物理基础模型铺路
未来巡天将产生数十亿张星系图像,但光谱数据相对稀少,亟需学习跨模态表征的模型。我们构建了包含134,533个星系图像(HSC-PDR2)和光谱(DESI-DR1)的数据集,采用多模态掩码自编码器(MMAE)将两者嵌入统一表征空间。MMAE基于变换器结构,通过掩码75%的数据并重建缺失的图像与光谱标记进行训练。实验验证了三种应用:在严重遮蔽下重建光谱与图像,以及仅从图像预测红移。模型可有效恢复星系形态、原子发射线峰值及宽频连续谱斜率等关键物理特征,但在细节和线强方面表现有限。在红移回归任务中,即便测试时无光谱,其预测散度仍优于或相当主流多模态模型。结果揭示了掩码自编码器在天体物理中的潜力与局限,推动向文本等更多模态扩展,以构建天体物理基础模型。
原文摘要 · Abstract (English)
Upcoming surveys will produce billions of galaxy images but comparatively few spectra, motivating models that learn cross-modal representations. We build a dataset of 134,533 galaxy images (HSC-PDR2) and spectra (DESI-DR1) and adapt a Multi-Modal Masked Autoencoder (MMAE) to embed both images and spectra in a shared representation. The MMAE is a transformer-based architecture, which we train by masking 75% of the data and reconstructing missing image and spectral tokens. We use this model to test three applications: spectral and image reconstruction from heavily masked data and redshift regression from images alone. It recovers key physical features, such as galaxy shapes, atomic emission line peaks, and broad continuum slopes, though it struggles with fine image details and line strengths. For redshift regression, the MMAE performs comparably or better than prior multi-modal models in terms of prediction scatter even when missing spectra in testing. These results highlight both the potential and limitations of masked autoencoders in astrophysics and motivate extensions to additional modalities, such as text, for foundation models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。