构建首个带文档上下文的科学示意图数据集,支持精准检索与分析。
DiagramBank: A Quality-Audited Dataset of Scientific Schematic Diagrams with Multi-Level Document Context

- 从AI/ML会议论文中提取5.7万张示意图,保留标题、摘要、图注等多层上下文。
- 经人工盲审验证,数据精度达93.67%,可平衡精度与覆盖范围。
- 适合从事科学文档理解、图谱检索与基准构建的研究者使用。
科学论文常使用示意图表达方法、流程与系统结构,但现有科学图像语料库常混入图表、截图和照片,且很少保留文档上下文。我们推出DiagramBank,一个从OpenReview主办的AI/ML会议中筛选出的57,100张示意图高质量数据集。每条记录关联图象与其论文标题、摘要、图注、文中引用片段、会议/年份元数据、来源字段及过滤标签。DiagramBank可支持科学文档理解、示意图检索、语料分析及未来基准建设。本文描述其提取与级联过滤流程、发布模式、置信度可控视图、数据卡与索引工具。人工盲审显示释放的级联过滤记录精度为93.67%;另通过CLIP阈值分析刻画了简单过滤视图的精度-覆盖率权衡。此外提供轻量级元数据索引与创作示例,展示下游应用协议,不将其视为独立方法。代码已开源:https://github.com/csml-rpi/DiagramBank。
原文摘要 · Abstract (English)
Scientific papers use schematic diagrams to communicate methods, workflows, and system structure, yet existing scientific-figure corpora often mix them with plots, screenshots, and photographs and rarely preserve document context. We introduce DiagramBank, a quality-audited dataset of 57,100 schematic diagrams curated from OpenReview-hosted AI/ML venues. Each record links a diagram image to its paper title, abstract, figure caption, in-text figure-reference spans, venue/year metadata, provenance fields, and filtering labels. DiagramBank is a reusable resource for scientific-document understanding, diagram retrieval, corpus analysis, and future benchmark construction. We describe its extraction and cascade-filtering pipeline, release schema, confidence-controlled views, dataset card, and indexing utilities. A manual blind audit of the released cascade-filtered records estimates 93.67% precision, and a separate CLIP threshold analysis characterizes the precision--coverage trade-off for simpler filtering views. We further provide lightweight metadata-indexing and authoring examples to illustrate downstream protocols without treating these utilities as standalone methods. The code is public at: https://github.com/csml-rpi/DiagramBank.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。