arXiv:2510.05016astro-ph.IMcs.AI2025-10被引 3

大模型在天文奥赛中表现接近人类顶尖水平,但推理能力仍有明显短板。

Large Language Models Achieve Gold Medal Performance at the International Olympiad on Astronomy & Astrophysics (IOAA)

  • 用国际天文奥赛试题系统评测五款大模型,涵盖多步推导与多模态分析。
  • 顶级模型得分超85%,达到金牌水平,理论题排名前二,数据题表现分化。
  • 空间想象与几何推理是普遍弱项,约52%-79%准确率,限制其科研自动化应用。

尽管特定任务展示大语言模型(LLMs)在部分天文学研究中已有初步成效,但这些实验仅呈现了其能力的局部图景,难以全面评估解决真实天文学问题所需的能力。现有基准和评测多集中于简单问答,主要考察天文知识,而忽视了实际研究中所需的复杂推理能力。为此,本文系统地在国际天文奥林匹克竞赛(IOAA)试题上评测五款先进大模型,这些试题旨在检验深层概念理解、多步推导及多模态分析能力。结果显示,Gemini 2.5 Pro与GPT-5平均得分分别为85.6%和84.2%,不仅达到金牌水平,且在2022–2025年四场理论考试中均位列约200–300名参赛者中的前两名。数据题方面,性能差异显著:GPT-5平均得分88.5%,排名前10;其他模型则降至48%–76%。深入错误分析表明,所有模型在概念推理、几何推理与空间可视化方面均存在持续弱点,准确率仅为52%–79%。因此,尽管大模型在理论题上逼近人类顶尖表现,但在成为自主科研代理前仍需弥补关键能力缺口。

原文摘要 · Abstract (English)

While task-specific demonstrations show early success in applying large language models (LLMs) to automate some astronomical research tasks, they only provide incomplete views of all necessary capabilities in solving astronomy problems, calling for more thorough understanding of LLMs' strengths and limitations. So far, existing benchmarks and evaluations focus on simple question-answering that primarily tests astronomical knowledge and fails to evaluate the complex reasoning required for real-world research in the discipline. Here, we address this gap by systematically benchmarking five state-of-the-art LLMs on the International Olympiad on Astronomy and Astrophysics (IOAA) exams, which are designed to examine deep conceptual understanding, multi-step derivations, and multimodal analysis. With average scores of 85.6% and 84.2%, Gemini 2.5 Pro and GPT-5 (the two top-performing models) not only achieve gold medal level performance but also rank in the top two among ~200-300 participants in all four IOAA theory exams evaluated (2022-2025). In comparison, results on the data analysis exams show more divergence. GPT-5 still excels in the exams with an 88.5% average score, ranking top 10 among the participants in the four most recent IOAAs, while other models' performances drop to 48-76%. Furthermore, our in-depth error analysis underscores conceptual reasoning, geometric reasoning, and spatial visualization (52-79% accuracy) as consistent weaknesses among all LLMs. Hence, although LLMs approach peak human performance in theory exams, critical gaps must be addressed before they can serve as autonomous research agents in astronomy.

大模型天文奥赛推理能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。