让天文图像支持自然语言对话,用大模型理解星空画面。
AstroLLaVA: towards the unification of astronomical data and natural language
- 在3万张天文图上微调,实现图文交互问答
- 支持开放问题,准确理解星体与现象的视觉信息
- 开源模型与数据集,推动天文AI协作
我们提出AstroLLaVA,一种面向天文学的视觉语言模型,可通过对天文图像进行自然语言对话实现交互。通过在约3万张图像上微调LLaVA模型,这些图像来自NASA‘每日一图’、欧洲南方天文台及NASA/ESA哈勃空间望远镜,包含图文描述和问答对。采用两阶段微调策略,使模型具备天文学领域的图像描述与视觉问答能力。我们在天文视觉问答基准上验证了性能,并公开模型权重、代码与训练数据集,以促进该领域开源研究。最后,我们提出一条通往通用天文数据与预训练语言模型对齐的路线图,为感兴趣的研究者提供开放协作平台。
原文摘要 · Abstract (English)
We present AstroLLaVA, a vision language model for astronomy that enables interaction with astronomical imagery through natural dialogue. By fine-tuning the LLaVA model on a diverse dataset of $\sim$30k images with captions and question-answer pairs sourced from NASA's `Astronomy Picture of the Day', the European Southern Observatory, and the NASA/ESA Hubble Space Telescope, we create a model capable of answering open-ended questions about astronomical concepts depicted visually. Our two-stage fine-tuning process adapts the model to both image captioning and visual question answering in the astronomy domain. We demonstrate AstroLLaVA's performance on an astronomical visual question answering benchmark and release the model weights, code, and training set to encourage further open source work in this space. Finally, we suggest a roadmap towards general astronomical data alignment with pre-trained language models, and provide an open space for collaboration towards this end for interested researchers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。