构建图文混杂论文的多模态信息提取基准,助力材料科学数据挖掘
MatViX: Multimodal Information Extraction from Visually Rich Articles
- 构建324篇论文与1688个结构化JSON数据集,覆盖文本、图表复杂关联
- 零样本评估显示现有模型在曲线相似性与层级对齐上表现不佳
- 提出专用模型DePlot提升曲线提取效果,适合材料领域研究者使用
多模态信息提取(MIE)对科学文献至关重要,因关键数据常分散于文字、图表和表格中。在材料科学领域,从研究论文中提取结构化信息可加速新材料发现。然而,科学内容的多模态特性与复杂关联性给传统文本方法带来挑战。我们提出 extsc{MatViX},一个由领域专家精心构建的基准,包含324篇完整研究论文和1688个复杂结构化的JSON文件,数据源自文档中的文本、表格与图表,提供全面的MIE挑战。我们引入评估方法,用于衡量曲线相似性准确度与层级结构对齐程度。同时,在零样本条件下对视觉语言模型(VLMs)进行基准测试,这些模型具备长上下文与多模态输入处理能力,并发现使用专用模型DePlot能显著提升曲线提取性能。结果表明当前模型仍有巨大改进空间。数据集与评估代码已公开。
原文摘要 · Abstract (English)
Multimodal information extraction (MIE) is crucial for scientific literature, where valuable data is often spread across text, figures, and tables. In materials science, extracting structured information from research articles can accelerate the discovery of new materials. However, the multimodal nature and complex interconnections of scientific content present challenges for traditional text-based methods. We introduce \textsc{MatViX}, a benchmark consisting of $324$ full-length research articles and $1,688$ complex structured JSON files, carefully curated by domain experts. These JSON files are extracted from text, tables, and figures in full-length documents, providing a comprehensive challenge for MIE. We introduce an evaluation method to assess the accuracy of curve similarity and the alignment of hierarchical structures. Additionally, we benchmark vision-language models (VLMs) in a zero-shot manner, capable of processing long contexts and multimodal inputs, and show that using a specialized model (DePlot) can improve performance in extracting curves. Our results demonstrate significant room for improvement in current models. Our dataset and evaluation code are available\footnote{\url{https://matvix-bench.github.io/}}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。