不同天文模型训练后,对星系物理的表征趋于一致。
The Platonic Universe: Do Foundation Models See the Same Sky?
- 用多种模型对比星系图像与光谱数据的表征相似性。
- 模型越大,表征越接近,表明存在共同宇宙认知模式。
- 适合关注跨模态统一表征的天文学与机器学习研究者。
我们通过测量在不同数据类型上训练的多种基础模型间的表征收敛性,检验天文学中的柏拉图表征假说(PRH)。利用詹姆斯·韦布空间望远镜(JWST)、HSC、Legacy Survey和DESI的光谱与成像观测数据,通过互近邻分析比较视觉变换器、自监督模型及天文学专用架构的表征。结果显示:随着模型容量提升,表征对齐程度普遍增强,支持向星系天体物理的共享表征收敛。结果表明,天文基础模型可采用预训练的通用架构,从而利用机器学习领域已投入的计算资源。
原文摘要 · Abstract (English)
We test the Platonic Representation Hypothesis (PRH) in astronomy by measuring representational convergence across a range of foundation models trained on different data types. Using spectroscopic and imaging observations from JWST, HSC, Legacy Survey, and DESI, we compare representations from vision transformers, self-supervised models, and astronomy-specific architectures via mutual $k$-nearest neighbour analysis. We observe consistent scaling: representational alignment generally increases with model capacity across our tested architectures, supporting convergence toward a shared representation of galaxy astrophysics. Our results suggest that astronomical foundation models can use pre-trained general-purpose architectures, allowing us to capitalise on the broader machine learning community's already-spent computational investment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。