构建科学图像理解基准,推动AI读懂论文中的图表与数据。
A Pathway to General-Purpose Scientific AI: Multimodal Comprehension of Scientific Images

- 构建包含1951张图表的ALD/E-ImageMiner基准,支持分类、表格提取等任务。
- 通过专家标注的视觉问答和摘要任务,评估AI对科学证据的理解能力。
- 面向未来科学推理,涵盖假设验证、溯源、不确定性等高级能力挑战。
科学图表和表格蕴含关键实验证据,但数字图书馆和多模态AI系统难以检索与解读。本文提出的ALD/E-ImageMiner基准及ICDAR 2026竞赛,包含来自205篇论文的1951张科学图表,经专家标注,涵盖分类、数据表提取、摘要生成与视觉问答等任务。这些任务旨在评估模型从视觉和定量层面理解科学内容的能力,包括领域相关的推理与证据支持。基于Bloom分类学设计的问题可促进更深层次的科学理解。本文提出以“从图像中获取科学概念理解”为长期目标,未来方向包括拓展至更广领域与图类型、上下文与跨文档综合、假设评估、溯源、不确定性建模、反事实推理及开放式的多模态科研。该视角将ICDAR 2026挑战与可行动的科学视觉知识、可验证的多模态科学AI愿景相连接。
原文摘要 · Abstract (English)
Scientific figures and tables encode essential experimental evidence, yet remain difficult for digital libraries and multimodal AI systems to retrieve and interpret. The ALD/E-ImageMiner benchmark and ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching Scientific Figures provide 1,951 figures from 205 publications, expert-annotated for classification, data table extraction, summarization, and visual question answering. In these companion proceedings, we present a forward-looking perspective on how the benchmark can guide future scientific-image challenges. We examine how its tasks probe capabilities from visual and quantitative reading to domain-grounded reasoning and evidential justification, and how Bloom-informed question design can support deeper scientific understanding. We propose "scientific conceptual understanding from images" as a long-term benchmark objective, with future directions including broader domains and figure types, contextual and cross-document synthesis, hypothesis evaluation, provenance, uncertainty, counterfactual grounding, and open-ended multimodal research. This perspective connects the ICDAR 2026 challenge to a broader agenda for machine-actionable scientific visual knowledge and verifiable multimodal scientific AI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。