arXiv:2606.10194cs.LGcs.AI2026-06被引 2

构建大规模多模态气候数据集,助力AI理解文本、图像与视频中的气候信息。

MMClima: A Framework for Multimodal Climate Science Data and Evaluation

论文配图:MMClima: A Framework for Multimodal Climate Science Data and Evaluation
图 1 · 摘自论文原文
  • 自动提取+人工验证,生成超10万条专家标注的多模态问答对。
  • 在跨模态推理任务中,新模型在文本问答上超越主流开源与闭源模型。
  • 适合气候科学与多模态AI研究者用于标准评估与模型训练。

气候变化研究日益依赖能跨文本、动态视觉内容和科学图表进行推理的AI系统,但现有气候问答基准规模小、以文本为主,涵盖模型有限。我们提出MMClima,一个大规模多模态气候问答框架,包含104,000+条专家验证的问答对,覆盖文章、视频字幕和图表,横跨五大核心气候科学领域。该框架通过自动化主张提取与问答合成,并结合人机协同验证,兼顾规模与可靠性。利用MMClima,我们对先进多模态语言模型在事实回忆、视觉解析和跨模态融合任务上的表现进行了评估。此外,在文本子集上微调得到mmclima-70b-txt,其在文本问答任务上优于多个强开源与闭源模型。我们发布数据集、评估流程、微调模型权重及数据构建框架,推动气候科学领域多模态评估的标准化。

原文摘要 · Abstract (English)

Climate change research increasingly requires AI systems that reason across text, dynamic visual content, and scientific figures, yet existing climate QA benchmarks are small, mostly textual, and cover a narrow range of models. We introduce MMClima, a large-scale multimodal climate question answering framework with 104k+ expert-validated question-answer pairs spanning articles, video transcriptions, and figures across five core climate science domains. MMClima is constructed via automated claim extraction and QA synthesis with human-in-the-loop validation to ensure both scale and reliability. Using MMClima, we benchmark state-of-the-art multimodal language models on tasks requiring factual recall, visual interpretation, and cross-modal synthesis. We additionally fine-tune on the textual split to produce mmclima-70b-txt, a domain-adapted baseline that outperforms strong open- and closed-source models on textual QA. We release the dataset, evaluation pipeline, fine-tuned model weights, and data creation framework to support standardized multimodal evaluation for climate science.

多模态气候科学问答系统数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。