用自监督学习统一天文光谱多分辨率数据,构建通用表示。
Universal Spectral Tokenization via Self-Supervised Panchromatic Representation Learning
- 自监督学习融合多源光谱数据,直接在原始波长网格上处理。
- 单模型实现跨分辨率与跨波段光谱统一表征,下游任务表现优异。
- 可作为天文学基础模型基石,或推广至气候、医疗等科学领域。
序列型科学数据涵盖多种分辨率与领域,将其统一为共同表示是构建科学领域基础模型的关键步骤。天文光谱是这一挑战的典型:大规模巡天已收集数百万条覆盖广泛波长和分辨率的光谱,但分析仍分散于不同光谱域(如光学与红外)和天体类型(如恒星与星系),限制了跨数据集的信息整合。我们提出一种深度学习模型,以自监督方式联合学习异构光谱数据。该通用光谱分词器可直接在原始波长网格上处理多种天体类型与分辨率的光谱,生成内在对齐、同质且物理意义明确的表示,能高效适配各类下游任务并取得竞争性性能。首次证明单一模型可统一跨分辨率与跨域的光谱数据,表明该模型可成为天文学基础模型的强大构建模块,并可能推广至气候、医疗等具有异构序列数据的科学领域。
原文摘要 · Abstract (English)
Sequential scientific data span many resolutions and domains, and unifying them into a common representation is a key step toward developing foundation models for the sciences. Astronomical spectra exemplify this challenge: massive surveys have collected millions of spectra across a wide range of wavelengths and resolutions, yet analyses remain fragmented across spectral domains (e.g., optical vs. infrared) and object types (e.g., stars vs. galaxies), limiting the ability to pool information across datasets. We present a deep learning model that jointly learns from heterogeneous spectra in a self-supervised manner. Our universal spectral tokenizer processes spectra from a variety of object types and resolutions directly on their native wavelength grids, producing intrinsically aligned, homogeneous, and physically meaningful representations that can be efficiently adapted to achieve competitive performance across a range of downstream tasks. For the first time, we demonstrate that a single model can unify spectral data across resolutions and domains, suggesting that our model can serve as a powerful building block for foundation models in astronomy -- and potentially extend to other scientific domains with heterogeneous sequential data, such as climate and healthcare.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。