专为太阳物理教育设计的问答模型,能准确解释复杂空间天气概念。
SolarGPT-QA: A Domain-Adaptive Large Language Model for Educational Question Answering in Space Weather and Heliophysics
- 基于LLaMA-3微调,融合科学文献与GPT-4生成的问答数据
- 在零样本下优于通用大模型,教育解释能力媲美指令微调模型
- 强调可读性与教学效果,适合科普与教学场景使用
太阳活动(如耀斑、日冕物质抛射和地磁暴)可能对卫星、航空、电网、数据中心及太空任务造成重大影响。极端事件常导致严重经济损失且预警时间短,凸显早期预警、精准预测与有效科普的重要性。尽管大语言模型在通用任务中表现良好,但缺乏领域知识与清晰解释复杂科学概念的教学能力。我们提出SolarGPT-QA,一个基于LLaMA-3的领域自适应问答系统,通过科学文献与大规模由GPT-4生成并经Grok-3优化的问答数据进行训练,采用学生友好的叙事风格。为评估回答质量,采用“大模型作为评判者”的评估框架,参考模型依据科学准确性、清晰度、完整性与教学有效性等结构化标准打分。结果表明,SolarGPT-QA在零样本设置下显著优于通用模型,在空间天气与日球物理学教育解释任务中达到与指令微调模型相当的性能。消融实验显示,领域自适应预训练与微调结合对平衡科学准确性与教学有效性至关重要。
原文摘要 · Abstract (English)
Solar activity, including solar flares, coronal mass ejections (CMEs), and geomagnetic storms can significantly impact satellites, aviation, power grids, data centers, and space missions. Extreme solar events can cause substantial economic damage with limited advance warning, underscoring the importance of early warning systems, accurate forecasting, and effective education in space science. Although large language models (LLMs) perform well on general tasks, they often lack domain specific knowledge and pedagogical capability to clearly explain complex space science concepts. We introduce SolarGPT-QA, a question answering system based on a domain adapted large language model built on the LLaMA-3 base model. The model is trained using scientific literature and large scale question and answer data generated with GPT-4 and refined using Grok-3 in a student friendly storytelling style. To evaluate response quality, we employ an LLM-as-judge evaluation framework, where a strong reference model assesses generated answers using structured criteria including scientific accuracy, clarity, completeness, and pedagogical effectiveness. Results show that SolarGPT-QA performs strongly relative to general purpose models in zero shot settings and achieves competitive performance compared to instruction tuned models for educational explanations in space weather and heliophysics. Ablation studies indicate that combining domain adaptive pretraining with fine tuning is important for balancing scientific accuracy and educational effectiveness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。