arXiv:2505.17592astro-ph.IMcs.LG2025-05被引 5

700亿参数天文专用模型,问答准确率达89%,媲美顶尖大模型。

AstroMLab 4: Benchmark-Topping Performance in Astronomy Q&A with a 70B-Parameter Domain-Specialized Reasoning Model

  • 基于700亿参数模型,用天文文献持续预训练并融合推理链。
  • 在3846道天文题上达到89.0%准确率,超越多数通用模型。
  • 适合天文学科研与教学,成本更低且专精度更高。

通用大语言模型虽具备广泛能力,但在专业领域知识上表现不足,限制其在天文学等高要求领域的可靠应用。本文基于Meta-Llama-3.1-70B基础模型,推出700亿参数的领域专用自然语言智能助手AstroSage-Llama-3.1-70B,面向天文学、天体物理、空间科学、宇宙学及天文仪器等领域。该模型通过海量天文文献进行持续预训练(CPT),再经监督微调(SFT)与模型融合优化,并在SFT数据集中引入推理链,使模型可选择直接回答或先输出人类可读的思考过程。在来自AstroMLab-1基准(Ting et al., 2024)的3846道验证问题上(训练时未使用相关文献),其准确率达89.0%,与GPT-5.2、Claude-4.5-Opus和Gemini-3-Pro相当,但更具成本优势。结果表明,大规模模型的领域专业化可使其在特定知识领域超越通用模型,推动人工智能在天文学中的前沿应用。

原文摘要 · Abstract (English)

General-purpose large language models (LLMs), despite their broad capabilities, often struggle with specialized domain knowledge. This gap hinders their deployment as reliable research agents in demanding fields such as astronomy. Building on our prior work with AstroSage-Llama-3.1-8B, this study introduces AstroSage-Llama-3.1-70B, a 70-billion parameter domain-specialized natural-language AI assistant. It is designed for research and education across astronomy, astrophysics, space science, astroparticle physics, cosmology, and astronomical instrumentation. Developed from the Meta-Llama-3.1-70B foundation, AstroSage-Llama-3.1-70B underwent extensive continued pre-training (CPT) on a vast corpus of astronomical literature, followed by supervised fine-tuning (SFT) and model merging. We integrated reasoning chains into the SFT dataset, enabling AstroSage-Llama-3.1-70B to either answer the user query immediately, or first emit a human-readable thought process. Evaluated on a validated subset of 3,846 questions from the AstroMLab-1 benchmark (Ting et al., 2024) -- derived from literature withheld during training -- AstroSage-Llama-3.1-70B achieves top-tier performance (89.0%), matching GPT-5.2, Claude-4.5-Opus, and Gemini-3-Pro while being more cost-efficient. This work demonstrates that domain specialization, when applied to large-scale models, can enable them to outperform generalist counterparts in specialized knowledge areas like astronomy, thereby advancing the frontier of AI capabilities in the field.

天文AI大模型推理链领域专用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。