首个天文科学计算与可视化评测基准,检验大模型真实科研辅助能力。
AstroVisBench: A Code Benchmark for Scientific Computing and Visualization in Astronomy
- 构建天文领域专用工作流生成与复杂绘图评测体系
- 验证5位天文学家标注后,用大模型作裁判的评估方法可靠
- 发现顶尖大模型在科研辅助上仍有显著能力鸿沟
大型语言模型(LLMs)正被探索用于科学研宄,包括文献综述、问题解答、研究思路生成甚至计算实验。最终目标是帮助科学家获得新洞见。在诸多科学领域,这些洞见常源于对数据的处理与可视化分析。然而,评估大模型驱动的科研流程是否输出了正确的科学认知仍具挑战,且此前研究未解决此问题。本文提出AstroVisBench,首个面向天文学领域的科学计算与可视化基准。该基准评估大模型能否(1)生成特定于天文学的数据处理与分析工作流,(2)通过复杂图表可视化结果。我们采用创新的“大模型作为裁判”评估流程,并经五位专业天文学家标注验证其有效性。基于AstroVisBench,我们对当前最先进语言模型进行评估,揭示其在作为科研助手方面存在显著能力差距。该评估为人工智能科学家提供了端到端的评测路径,推动以可视化为核心的科研工作流发展,适用于物理、生物等多个领域。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are being explored for applications in scientific research, including their capabilities to synthesize literature, answer research questions, generate research ideas, and even conduct computational experiments. Ultimately, our goal is for these to help scientists derive novel scientific insights. In many areas of science, such insights often arise from processing and visualizing data to understand its patterns. However, evaluating whether an LLM-mediated scientific workflow produces outputs conveying the correct scientific insights is challenging to evaluate and has not been addressed in past work. We introduce AstroVisBench, the first benchmark for both scientific computing and visualization in the astronomy domain. AstroVisBench judges a language model's ability to both (1) create astronomy-specific workflows to process and analyze data and (2) visualize the results of these workflows through complex plots. Our evaluation of visualizations uses a novel LLM-as-a-judge workflow, which is validated against annotation by five professional astronomers. Using AstroVisBench we present an evaluation of state-of-the-art language models, showing a significant gap in their ability to engage in astronomy research as useful assistants. This evaluation provides a strong end-to-end evaluation for AI scientists that offers a path forward for the development of visualization-based workflows, which are central to a broad range of domains from physics to biology.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。