arXiv:2607.26848cs.CV2026-07被引 2

评测多模态AI在原子层沉积/刻蚀科学图示中的信息抽取能力

ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures

论文配图:ICDAR 2026 Competition on Information Extraction from Atomic Layer Deposition/Etching (ALD/E) Scientific Figures
图 1 · 摘自论文原文
  • 构建专家标注的多任务基准数据集Sci-ImageMiner
  • 顶尖模型在分类/摘要上表现好,但数据提取与推理差
  • 适合研究科学图像理解与多模态推理的学者参考

利用多模态AI理解科学图表并进行推理,需融合视觉感知与领域知识,以提取文本未明确表达的信息。本次竞赛配套的Sci-ImageMiner基准数据集,通过跨四个端到端互补任务的专家标注,提升了以往科学类竞赛的标准。比赛从2026年1月9日至4月8日,吸引68名活跃参赛者和1,263次公开/私密提交。结果表明,当前最先进的多模态模型在分类与摘要任务中表现良好,但在数据抽取与科学推理(尤其是视觉问答)方面仍存在显著不足。该发现揭示了现有系统的关键局限,指明了提升领域感知型多模态AI的挑战与机遇。整体而言,Sci-ImageMiner基准与竞赛为推进科学图表理解与推理研究提供了严谨平台,展示了先进方法在这一复杂研究领域的潜力。

原文摘要 · Abstract (English)

Scientific figure comprehension and reasoning using multimodal AI requires integrating visual perception with domain-specific reasoning to extract meaningful knowledge, often not presented in the text of a research publication. The Sci-ImageMiner benchmark dataset, accompanied by a community-driven competition, raises the bar over prior scientific competitions by curating a comprehensive, expert-annotated dataset across four end-to-end complementary tasks. The competition attracted 68 active participants and 1,263 public/private submissions from 9th January 2026 to 8th April 2026. Our results show that state-of-the-art multimodal models perform well on classification and summarization tasks but struggle with data extraction and scientific reasoning, particularly in visual question-answering. These findings reveal key limitations and highlight challenges and opportunities for improving domain-aware multimodal AI systems. Overall, the Sci-ImageMiner benchmark and competition establish a rigorous platform for advancing research in scientific figure comprehension and reasoning and demonstrate the potential of state-of-the-art approaches for a challenging and complex research area.

科学图示多模态AI信息抽取视觉问答

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。