评测大模型在天文事件分类中的准确率、推理能力和诚实度。
AstroAlertBench: Evaluating the Accuracy, Reasoning, and Honesty of Multimodal LLMs in Astronomical Classification

- 构建三阶段逻辑链评估框架:元数据定位、科学推理、五类分类。
- 基于1500条真实警报测试13个模型,发现高准确率未必代表可信。
- 强调模型自我评估能力,适合天文学家和AI可解释性研究者。
现代天文观测台产生海量多模态数据,人工专家审核成为关键瓶颈。尽管多模态大语言模型(LLMs)在解析复杂视觉与文本输入方面展现潜力,其在专业科学分类中提供可解释推理的能力仍缺乏系统评估。我们提出AstroAlertBench,一个全面的多模态基准,用于评估LLM在天文事件审核中的表现,涵盖三个阶段的逻辑链条:元数据定位、科学推理与五类层次化分类。基于来自宽视场巡天望远镜(Zwicky Transient Facility, ZTF)的1500条真实警报样本进行测试,我们评估了13个支持视觉输入的前沿闭源与开源模型。结果表明,高准确率并不总与模型“诚实度”一致——即自我评估推理能力,这影响其作为实际助手的可靠性。我们还建立人机协同评估协议,为未来社区规模参与奠基。AstroAlertBench为开发校准且可解释的天文智能助手提供了框架。
原文摘要 · Abstract (English)
Modern astronomical observatories generate a massive volume of multimodal data, creating a critical bottleneck for expert human review. While multimodal large language models (LLMs) have shown promise in interpreting complex visual and textual inputs, their ability to perform specialized scientific classification while providing interpretable reasoning remains understudied. We introduce AstroAlertBench, a comprehensive multimodal benchmark designed to evaluate LLM performance in astronomical event review along a three-stage logical chain: metadata grounding, scientific reasoning, and hierarchical classification over five categories. We use a pilot sample of 1,500 real-world alerts from the Zwicky Transient Facility (ZTF), a wide-field survey that scans the northern sky to detect transient astronomical events. On this dataset, we benchmark 13 frontier closed-source and open-weight LLMs that support visual input. Our results reveal that high accuracy does not always align with model ``honesty,'' defined as the ability to self-evaluate its reasoning, which affects its reliability as a real-world assistant. We further initialize a human-in-the-loop evaluation protocol as a precursor to future community-scale participation. Together, AstroAlertBench provides a framework for developing calibrated and interpretable astronomical assistants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。