GPT-4V生成的科学图表标题被编辑首选,显著优于其他模型。
Do Large Multimodal Models Solve Caption Generation for Scientific Figures? Lessons Learned from SciCap Challenge 2023
- 对比多种大模型在科学图表字幕生成上的表现
- GPT-4V生成结果获专业编辑最高青睐
- 揭示当前大模型在科研图文理解中的突破
自2021年SciCap数据集发布以来,学术界在生成学术文章中科学图表字幕方面取得了显著进展。2023年首届SciCap挑战赛启动,邀请全球团队利用扩展的SciCap数据集,为跨学科、多类型科学图表开发字幕生成模型。与此同时,文本生成模型快速发展,涌现出多个强大预训练的大规模多模态模型(LMMs),在视觉-语言任务中表现出色。本文综述了首届SciCap挑战赛,并详细分析了各类模型在该数据集上的表现,呈现了该领域的现状。研究发现,专业编辑对GPT-4V生成的图表字幕评价最高,甚至优于其他所有模型及作者原稿。基于此关键发现,我们深入分析了先进LMM是否已解决科学图表字幕生成任务。
原文摘要 · Abstract (English)
Since the SciCap datasets launch in 2021, the research community has made significant progress in generating captions for scientific figures in scholarly articles. In 2023, the first SciCap Challenge took place, inviting global teams to use an expanded SciCap dataset to develop models for captioning diverse figure types across various academic fields. At the same time, text generation models advanced quickly, with many powerful pre-trained large multimodal models (LMMs) emerging that showed impressive capabilities in various vision-and-language tasks. This paper presents an overview of the first SciCap Challenge and details the performance of various models on its data, capturing a snapshot of the fields state. We found that professional editors overwhelmingly preferred figure captions generated by GPT-4V over those from all other models and even the original captions written by authors. Following this key finding, we conducted detailed analyses to answer this question: Have advanced LMMs solved the task of generating captions for scientific figures?
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。