arXiv:2605.22162astro-ph.IMastro-ph.SR2026-05

用大语言模型分析恒星光谱,实现高精度参数与元素含量推断

Spectra as Language: Large Language Models for Scalable Stellar Parameter and Abundance Inference

论文配图:Spectra as Language: Large Language Models for Scalable Stellar Parameter and Abundance Inference
图 1 · 摘自论文原文
  • 将语言模型思路迁移至恒星光谱,构建两阶段推断框架
  • 可准确估计有效温度、表面重力、金属量及约20种元素丰度
  • 数据越多效果越好,适合未来大规模光谱巡天应用

恒星光谱蕴含恒星物理性质与化学组成的关键信息。精确的恒星参数测定对解答星系与恒星演化等重大问题至关重要。大规模光谱巡天已积累前所未有的光谱数据。传统特征提取或模型拟合方法在高维、海量数据下面临泛化能力有限、计算效率低等问题。近期大语言模型在自然语言处理、DNA/RNA序列分析、蛋白质/化学分子解析等任务中展现出强大的泛化与特征学习能力。恒星光谱是连续的序列信号,具备向语言模型迁移的潜力。本文提出一种两阶段大语言模型框架,用于恒星参数推断,可准确估计有效温度、表面重力、金属量以及约20种化学元素的丰度。尺度律分析表明,随着数据量增加,性能系统性提升,为未来大规模光谱巡天提供了可扩展的解决方案。

原文摘要 · Abstract (English)

Stellar spectra encode key information on the physical properties and chemical compositions of stars. Accurate stellar parameter determination is essential for addressing major questions such as galaxy and stellar evolution. Large-scale spectroscopic surveys have accumulated unprecedented spectral data. Traditional feature extraction or model-fitting approaches struggle with high-dimensional, massive datasets, limited generalization, and computational inefficiency. Recent advances in large language models demonstrate strong generalization and feature-learning in tasks like natural language processing, DNA/RNA sequence analysis, and protein/chemical parsing. Stellar spectra are continuous sequential signals, enabling the transfer of language models to stellar spectroscopy. Here, we propose a two-stage large language model framework for stellar parameter inference, achieving accurate estimation of effective temperature, surface gravity, metallicity, and abundances of ~20 chemical elements. Scaling-law analyses show systematic performance improvements with increasing data, providing a scalable framework for forthcoming large-scale surveys.

恒星光谱大模型参数推断天体物理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。