arXiv:2510.23749astro-ph.IMcs.LG2025-10中稿 · NeurIPS被引 3

用稀疏自编码器发现银河形态中隐藏的天体物理特征。

Re-envisioning Euclid Galaxy Morphology: Identifying and Interpreting Features with Sparse Autoencoders

  • 通过稀疏自编码器从预训练模型中提取可解释的单一语义特征。
  • 在欧几里得数据上,特征与星系分类标签的对齐度优于主成分分析。
  • 适合天体物理学家探索人类分类体系之外的新现象。

稀疏自编码器(SAEs)能高效地从预训练神经网络中识别出候选的单义性特征,用于星系形态分析。我们在欧几里得项目第一阶段(Euclid Q1)图像上,分别使用监督学习(Zoobot)和新型自监督学习(MAE)模型进行了验证。公开发布的MAE模型实现了超人类水平的图像重建性能。虽然基于监督模型的主成分分析(PCA)主要捕捉到与星系动物园(Galaxy Zoo)决策树一致的特征,但SAEs能够发现超出该框架的可解释特征。此外,SAE特征与星系动物园标签的对齐度更强。尽管解释性仍存挑战,但SAEs为发现人类分类体系之外的天体物理现象提供了强大工具。

原文摘要 · Abstract (English)

Sparse Autoencoders (SAEs) can efficiently identify candidate monosemantic features from pretrained neural networks for galaxy morphology. We demonstrate this on Euclid Q1 images using both supervised (Zoobot) and new self-supervised (MAE) models. Our publicly released MAE achieves superhuman image reconstruction performance. While a Principal Component Analysis (PCA) on the supervised model primarily identifies features already aligned with the Galaxy Zoo decision tree, SAEs can identify interpretable features outside of this framework. SAE features also show stronger alignment than PCA with Galaxy Zoo labels. Although challenges in interpretability remain, SAEs provide a powerful engine for discovering astrophysical phenomena beyond the confines of human-defined classification.

星系形态稀疏自编码器自监督学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。