arXiv:2608.23626cs.AIastro-ph.IM2026-08综述

天文基础模型受探测通道干扰,导致红移估计严重偏差。

A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts

论文配图:A survey detection channel overrides the pixels in an astronomical foundation model, and biases tomographic mean redshifts
图 1 · 摘自论文原文
  • 通过因果干预发现,探测掩膜比图像像素主导模型输出
  • 红移误差达LSST标准的0.71倍,最大超8.3倍
  • 模型规模越大偏差越明显,谱数据编码更易被干扰

天文学基础模型在包含像素和星表产物的数据上训练,而这些星表存在可测量的不完整性。本文对训练了2亿多个天体的39模态Transformer AION-1进行因果干预:保持图像标记完全一致,仅修改探测分割图,使模型报告的所有量(通量、大小、椭圆率、红移)变化幅度达对照组的110至4400倍。关键机制是检测门控,位置中心性(r=0.47)影响远大于遮罩内光强(r=0.30);在322个真实混合源中,模型无视光分布划分方式(R=-0.006)。非特定通道偏好:矛盾星表光度使模型性能下降9倍,甚至不如无元数据。遗产巡天管道有3.68%目标未被分割覆盖,若以实际返回场次表示漏检,会导致40次分类中红移均值偏移达中位0.71倍,12次超出要求;观测位置误差最坏情况达8.3倍。按测光依赖画出漏检也不改变结果。光谱可消除影响,屏蔽探测通道亦可消除且无代价,且偏差随模型规模增大。两个限制存在于分词器:图像编码器对源块仅分辨28种有效状态,而光谱编码器可达934种;红移读出受量化限制。稀疏字典作为因果干预工具不可靠,15次恢复成功率仅26%-75%,种子单独变动可致18点偏移。

原文摘要 · Abstract (English)

Foundation models for astronomy are trained on survey pixels together with the catalogue products derived from those pixels. Those catalogues are incomplete at a measurable rate, and a model trained on both inherits that incompleteness as a systematic. We audit AION-1, a 39-modality transformer trained on more than 200 million objects, using causal interventions on its inputs. Holding the image tokens byte-identical and editing only the survey segmentation map changes every quantity the model reports -- flux, size, ellipticity, redshift -- by 110-4400 times a matched placebo. The mechanism is detection gating, presence at the field centre (r = 0.47), not the light the mask encloses (r = 0.30); across 322 real blends the model ignores how the pipeline partitioned the light (R = -0.006). Nor is the preference specific to that channel: contradicted catalogue photometry leaves the model nine times worse than supplying no metadata at all. The Legacy Survey pipeline leaves 3.68% of targets with no segment covering their position. Propagating that rate, with a miss represented by the fields the pipeline actually returns, shifts tomographic mean redshifts by a median 0.71 times the LSST DESC requirement over 40 assignments and exceeds it in 12; observed positional errors take the worst bin to 8.3 times. Drawing the misses by their measured magnitude dependence rather than uniformly does not change it. Spectroscopy removes the effect, withholding the detection channel removes it at no measurable cost, and the effect grows with model scale. Two further limits lie in the tokeniser: its image codec resolves 28 effective states on source patches against 934 for the spectrum codec, and the redshift readout is quantisation-limited. Sparse dictionaries are unreliable causal handles: across 15, recovery spans 26-75% and moves up to 18 points on the seed alone.

天文模型因果干预红移偏差基础模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。