用高质量天文书摘持续预训练,提升天文大模型性能。
AstroMLab 2: AstroLLaMA-2-70B Model and Benchmarking Specialised LLMs for Astronomy
- 用arXiv摘要数据持续预训练,缓解模型性能下降。
- 70B参数模型经预训练后表现显著提升,小模型则出现灾难性遗忘。
- 新推出AstroLLaMA-2-70B和AstroLLaMA-3-8B,助力天文AI研究。
在领域特定数据上持续预训练大语言模型已被提出以提升下游任务表现。然而,天文学领域此前缺乏专用评估基准,阻碍了专用大模型的客观评测。本研究基于近期整理的高质量天文多项选择题数据集,首次对天文专用大模型进行量化评估。结果发现,此前发布的基于LLaMA-2-7B的AstroLLaMA系列模型表现反而低于基础模型。我们证明,通过使用高质量数据(如arXiv摘要)进行持续预训练,可部分缓解这一性能下降问题。尽管小模型存在灾难性遗忘现象,但70B参数模型在持续预训练后仍实现显著提升。然而,当前监督微调数据集仍限制了指令型模型的表现。本研究同步发布了AstroLLaMA-3-8B与AstroLLaMA-2-70B两个新模型,延续前序系列。
原文摘要 · Abstract (English)
Continual pretraining of large language models on domain-specific data has been proposed to enhance performance on downstream tasks. In astronomy, the previous absence of astronomy-focused benchmarks has hindered objective evaluation of these specialized LLM models. Leveraging a recent initiative to curate high-quality astronomical MCQs, this study aims to quantitatively assess specialized LLMs in astronomy. We find that the previously released AstroLLaMA series, based on LLaMA-2-7B, underperforms compared to the base model. We demonstrate that this performance degradation can be partially mitigated by utilizing high-quality data for continual pretraining, such as summarized text from arXiv. Despite the observed catastrophic forgetting in smaller models, our results indicate that continual pretraining on the 70B model can yield significant improvements. However, the current supervised fine-tuning dataset still constrains the performance of instruct models. In conjunction with this study, we introduce a new set of models, AstroLLaMA-3-8B and AstroLLaMA-2-70B, building upon the previous AstroLLaMA series.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。