arXiv:2608.22614cs.LGastro-ph.IM2026-08中稿 · COLM

用星系图像训练模型,揭示了概念在神经网络中出现的顺序规律。

What AstroPT knows about galaxies, and what that can teach us about LLMs

论文配图:What AstroPT knows about galaxies, and what that can teach us about LLMs
图 1 · 摘自论文原文
  • 基于天文数据训练模型,利用真实物理关系验证概念涌现顺序。
  • 像素直接相关的量(如波段星等)早期且浅层可解码,复杂推断量(如红移)后期深层出现。
  • 结果与训练目标无关,适合用于校准大模型可解释性方法。

可解释性研究常关注概念何时在训练中出现以及线性探测是否能恢复真实结构,但语言模型缺乏可验证的概念排序和关系基准。我们提出使用天文领域的真实知识作为校准基准——通过数百万张星系图像训练的AstroPT模型。该模型具有已知的概念难度排序和物理关系结构。对不同训练检查点、层数、模型规模和目标函数下的冻结表征进行探测发现:直接从像素中读取的量(如波段星等)在训练初期、网络浅层即可被解码;而依赖多波段或光谱、需推断的量(如红移、恒星形成率)则更晚、更深地出现。此顺序在多种训练设置下保持不变,仅随模型容量增大而幅度变化,不改变序列。线性探测方向还恢复了星系属性间的已知物理结构。研究表明,天文学为校准无监督的大模型可解释性方法提供了可控实验环境。

原文摘要 · Abstract (English)

Interpretability research increasingly asks when concepts emerge during training and whether linear probes recover real structure, but in language models these claims are hard to validate because language offers little ground-truth ordering of concepts or relationships among them. We propose the use of astronomical ground truth through AstroPT, a transformer trained on millions of galaxy images, as a calibration testbed. AstroPT is an LLM-like model trained within a domain where the difficulty ordering of concepts and the relations among them are known in advance. Probing frozen representations across checkpoints, layers, model sizes, and objective choices, we find that galaxy properties emerge in a fixed order that tracks their known difficulty---quantities written almost directly into the pixels (band magnitude) become decodable early in training and shallow in the network, while multiband/spectra based and inferred quantities (such as redshift and specific star formation rate) emerge later and deeper. This order is invariant to our tested training objectives, and scales in magnitude but not in sequence with capacity. Our linear probe directions further recover the known physical structure among galaxy properties. Our findings suggest that astronomy offers a controlled sandbox for calibrating mechanistic interpretability methods we otherwise apply to LLMs blind.

可解释性星系生成模型探针机制分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。