构建大规模多模态气候数据集,助力AI理解文本、图像与视频中的气候信息。
MMClima: A Framework for Multimodal Climate Science Data and Evaluation

- 自动提取+人工验证,生成超10万条专家标注的多模态问答对。
- 在跨模态推理任务中,新模型在文本问答上超越主流开源与闭源模型。
- 适合气候科学与多模态AI研究者用于标准评估与模型训练。
气候变化研究日益依赖能跨文本、动态视觉内容和科学图表进行推理的AI系统,但现有气候问答基准规模小、以文本为主,涵盖模型有限。我们提出MMClima,一个大规模多模态气候问答框架,包含104,000+条专家验证的问答对,覆盖文章、视频字幕和图表,横跨五大核心气候科学领域。该框架通过自动化主张提取与问答合成,并结合人机协同验证,兼顾规模与可靠性。利用MMClima,我们对先进多模态语言模型在事实回忆、视觉解析和跨模态融合任务上的表现进行了评估。此外,在文本子集上微调得到mmclima-70b-txt,其在文本问答任务上优于多个强开源与闭源模型。我们发布数据集、评估流程、微调模型权重及数据构建框架,推动气候科学领域多模态评估的标准化。
原文摘要 · Abstract (English)
Climate change research increasingly requires AI systems that reason across text, dynamic visual content, and scientific figures, yet existing climate QA benchmarks are small, mostly textual, and cover a narrow range of models. We introduce MMClima, a large-scale multimodal climate question answering framework with 104k+ expert-validated question-answer pairs spanning articles, video transcriptions, and figures across five core climate science domains. MMClima is constructed via automated claim extraction and QA synthesis with human-in-the-loop validation to ensure both scale and reliability. Using MMClima, we benchmark state-of-the-art multimodal language models on tasks requiring factual recall, visual interpretation, and cross-modal synthesis. We additionally fine-tune on the textual split to produce mmclima-70b-txt, a domain-adapted baseline that outperforms strong open- and closed-source models on textual QA. We release the dataset, evaluation pipeline, fine-tuned model weights, and data creation framework to support standardized multimodal evaluation for climate science.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。