arXiv:2602.13758cs.CVcs.AI2026-02KDD被引 4

构建150万组科学图像数据集,提升大模型对科研图表的理解能力。

OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding

  • 用多模态大模型动态重生成高密度图文描述,融合视觉与文本信息。
  • 数据集覆盖10+学科,图文相似度从0.769提升至0.956,显著增强语义准确率。
  • 适合研究科学图像理解、多模态模型训练与评估的学者使用。

多模态大语言模型在自然图像理解上表现优异,但在解释科研图像(如示意图、实验表征图、分析图表)方面能力有限,尤其在开源模型中更为明显。这一差距主要源于现有数据集领域覆盖窄、结构标注粗略、语义基础薄弱。本文提出OmniScience,一个大规模、高保真多模态数据集,包含超过150万组图-文-上下文三元组,覆盖10多个主要科学领域。为生成信息密度更高、准确性更强的图像描述,我们设计动态模型路由重生成管道,利用先进多模态大模型联合融合视觉特征、原始图注及科学家撰写的文本引用,生成密集且自洽的描述。该管道结合严格的质量过滤与人类专家判断对齐,确保事实准确与语义完整,使图文多模态相似度从0.769提升至0.956。我们进一步提出一种基于问答的图理解评估协议作为代理任务。在该设置下,基于OmniScience微调的Qwen2.5-VL-3B模型相较基线显著提升,在MM-MT-Bench上增益0.378,在MMMU上增益0.140。

原文摘要 · Abstract (English)

Multimodal Large Language Models demonstrate strong performance on natural image understanding, yet exhibit limited capability in interpreting scientific images, including but not limited to schematic diagrams, experimental characterizations, and analytical charts. This limitation is particularly pronounced in open-source MLLMs. The gap largely stems from existing datasets with limited domain coverage, coarse structural annotations, and weak semantic grounding. We introduce OmniScience, a large-scale, high-fidelity multi-modal dataset comprising 1.5 million figure-caption-context triplets, spanning more than 10 major scientific disciplines. To obtain image caption data with higher information density and accuracy for multi-modal large-model training, we develop a dynamic model-routing re-captioning pipeline that leverages state-of-the-art multi-modal large language models to generate dense, self-contained descriptions by jointly synthesizing visual features, original figure captions, and corresponding in-text references authored by human scientists. The pipeline is further reinforced with rigorous quality filtering and alignment with human expert judgments, ensuring both factual accuracy and semantic completeness, and boosts the image-text multi-modal similarity score from 0.769 to 0.956. We further propose a caption QA protocol as a proxy task for evaluating visual understanding. Under this setting, Qwen2.5-VL-3B model finetuned on OmniScience show substantial gains over baselines, achieving a gain of 0.378 on MM-MT-Bench and a gain of 0.140 on MMMU.

科学图像多模态数据集大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。