arXiv:2606.03410cs.CV2026-06

首个工程图谱视觉语言理解数据集,评测AI读图能力。

Enginuity: A Dataset and Benchmark for Vision-Language Understanding of Engineering Diagrams

论文配图:Enginuity: A Dataset and Benchmark for Vision-Language Understanding of Engineering Diagrams
图 1 · 摘自论文原文
  • 构建工程图谱多任务评估框架,含零件表提取与问答
  • 模型识别零件准确率高但描述质量差,事实推理能力普遍不足
  • 提出语义评价方法,解决传统匹配指标低估真实能力问题

工程图谱对视觉语言模型构成独特挑战:其信息密集布局、领域专用符号及图文标注与结构化零件表之间的交叉引用,不同于自然图像或通用文档。尽管在维修、服务和设计流程中至关重要,该领域尚无公开基准。为此,我们推出Enginuity——首个面向复杂工程图谱的开放数据集与评估基准。基于美国军用维修手册,定义两项任务:结构化零件表提取(任务1)与自由形式图谱问答(任务2)。评估四种前沿视觉语言模型(GPT-5.2 Chat, Claude Opus 4.7, Gemma 4, Qwen3-VL-32B-Instruct)在零样本与思维链提示下的表现。任务1中,模型召回率达0.61–0.87,但标记词级F1仅为0.03–0.18,暴露识别与描述精度间的系统性差距;任务2显示所有模型均存在一致的事实推理缺陷。辅助分析表明,基于词重叠的指标相比语义相似度,低报模型能力2–6倍,推动采用大模型为裁判进行领域特定评估校准。数据集、标注、评估工具链及模型输出均已开源,支持可复现研究。

原文摘要 · Abstract (English)

Engineering diagrams pose a distinct challenge for vision-language models: unlike natural images or general documents, they encode information through dense spatial layouts, domain-specific symbols, and cross-references between visual callouts and structured parts tables. Despite their centrality to service, repair, and design workflows, there is no public benchmark for measuring VLM capabilities in this domain; existing datasets primarily focus on flowcharts, scientific figures, or business documents. To address this gap, we introduce Enginuity, the first open dataset and benchmark for evaluating VLMs on complex engineering diagrams. We define two tasks over a corpus of U.S. military service and repair manuals: structured parts-table extraction (Task 1) and free-form visual diagram question answering (VQA)(Task 2) for benchmarking. We evaluate four frontier VLMs (GPT-5.2 Chat, Claude Opus 4.7, Gemma 4, Qwen3-VL-32B-Instruct) under zero-shot and chain-of-thought prompting. On Task 1, models reach Recall@all of 0.61-0.87 but Token F1pen of only 0.03-0.18, exposing a systematic gap between part identification and description fidelity. Task 2 reveals a consistent factual-reasoning gap across all models. A supporting analysis shows that token-overlap metrics under-report model capability on technical descriptions by 2-6x relative to semantic similarity, motivating LLM-as-judge calibration for domain-specific evaluation. We release the dataset, annotations, evaluation harness, and per-sample model outputs to support a reproducible study of VLM capability on engineering content.

工程图谱视觉语言模型评测基准多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。