arXiv:2601.00264cs.CV2026-01

构建1550万对图文数据,提升科学图像与文本的理解能力

S1-MMAlign: A Large-Scale, Multi-Disciplinary Dataset for Scientific Figure-Text Understanding

  • 用大模型从摘要和引用中重构科学图注,解决图文对齐问题
  • 增强后数据使CLIP图文对齐提升,零样本科学图文生成性能显著提高
  • 适合研究AI for Science、科学多模态模型的学者使用

多模态学习在通用领域取得突破,但在科学发现中受限于复杂科学图像与简略文字描述之间的语义鸿沟。我们提出S1-MMAlign,一个涵盖超过1550万对高质量图文的跨学科大规模多模态数据集,源自250万篇开放获取的科学论文。覆盖物理、生物、工程等多领域,包含实验装置、热力图、显微图像等多种视觉模态。为解决原始图注普遍存在的弱对齐问题,我们设计了一套面向AI的语义增强流程,利用先进多模态大模型,结合论文摘要与对应图表的引用上下文,合成全面的图像描述。技术验证表明,该流程显著降低SciBERT伪困惑度,提升CLIP图文对齐效果,并大幅改善多模态大模型在零样本科学图注生成、跨领域科学推理和视觉指令微调任务中的表现。S1-MMAlign为人工智能驱动的科学理解提供了关键基础资源,支持科学基础模型开发及各类下游科学智能应用。数据集已公开:https://huggingface.co/datasets/ScienceOne-AI/S1-MMAlign。

原文摘要 · Abstract (English)

Multimodal learning has revolutionized general domain tasks, yet its application in scientific discovery is hindered by the profound semantic gap between complex scientific imagery and sparse textual descriptions. We present S1-MMAlign, a large-scale, multi-disciplinary multimodal dataset comprising over 15.5 million high-quality image-text pairs derived from 2.5 million open-access scientific papers. Spanning disciplines from physics and biology to engineering, the dataset captures diverse visual modalities including experimental setups, heatmaps, and microscopic imagery. To address the pervasive issue of weak alignment in raw scientific captions, we introduce an AI-ready semantic enhancement pipeline that leverages advanced multimodal large language models to recaption images, by synthesizing comprehensive context from paper abstracts and the citation contexts of corresponding figures. Technical validation confirms that our enhancement pipeline markedly improves data quality via reduced SciBERT pseudo-perplexity and enhanced CLIP image-text alignment, while also significantly boosting multimodal large language models performance in zero-shot scientific captioning, multi-domain scientific reasoning, and visual instruction tuning. S1-MMAlign provides a pivotal foundational resource for cross-modal scientific understanding in the AI for Science era, supporting the development of scientific foundation models and a wide range of downstream scientific intelligence applications. The dataset is publicly available at https://huggingface.co/datasets/ScienceOne-AI/S1-MMAlign.

科学多模态图文对齐数据集

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。