arXiv:2510.00063astro-ph.IMcs.AI2025-10被引 4

首个专为天文学设计的多模态大模型评测基准,填补领域空白。

AstroMMBench: A Benchmark for Evaluating Multimodal Large Language Models Capabilities in Astronomy

  • 构建涵盖6个子领域的621道多选题,由15位专家审核确保专业性。
  • 25个模型测试显示,Ovis2-34B准确率达70.5%最高,闭源模型仍占优。
  • 揭示模型在宇宙学等高难度领域表现差,适合天文学与AI交叉研究者使用。

天文学图像理解对多模态大语言模型(MLLMs)应用于专业科学任务构成重大挑战。现有基准侧重通用多模态能力,未能反映天文数据的复杂性。为此,我们提出AstroMMBench,首个专注于天文图像理解的综合性评测基准。该基准包含621道多项选择题,覆盖六个天体物理子领域,由15位领域专家共同审校以保证质量和相关性。我们对25种不同的MLLMs进行了广泛评估,其中包括22个开源和3个闭源模型。结果表明,Ovis2-34B取得最高整体准确率(70.5%),甚至优于部分强闭源模型。各模型在六个子领域的表现存在差异,在宇宙学和高能天体物理等领域尤为困难,而在仪器与太阳天体物理等领域相对较好。这些发现凸显了像AstroMMBench这样的领域专用基准在评估和引导MLLM针对科学应用发展的关键作用。AstroMMBench为人工智能与天文学交叉领域的进步提供了基础资源和动态工具。

原文摘要 · Abstract (English)

Astronomical image interpretation presents a significant challenge for applying multimodal large language models (MLLMs) to specialized scientific tasks. Existing benchmarks focus on general multimodal capabilities but fail to capture the complexity of astronomical data. To bridge this gap, we introduce AstroMMBench, the first comprehensive benchmark designed to evaluate MLLMs in astronomical image understanding. AstroMMBench comprises 621 multiple-choice questions across six astrophysical subfields, curated and reviewed by 15 domain experts for quality and relevance. We conducted an extensive evaluation of 25 diverse MLLMs, including 22 open-source and 3 closed-source models, using AstroMMBench. The results show that Ovis2-34B achieved the highest overall accuracy (70.5%), demonstrating leading capabilities even compared to strong closed-source models. Performance showed variations across the six astrophysical subfields, proving particularly challenging in domains like cosmology and high-energy astrophysics, while models performed relatively better in others, such as instrumentation and solar astrophysics. These findings underscore the vital role of domain-specific benchmarks like AstroMMBench in critically evaluating MLLM performance and guiding their targeted development for scientific applications. AstroMMBench provides a foundational resource and a dynamic tool to catalyze advancements at the intersection of AI and astronomy.

多模态天文学评测基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。