首个工业产品多图属性提取基准,揭示大模型跨图整合信息能力不足。
IndustryBench-MIPU: Benchmarking Multi-Image Attribute Value Extraction for Industrial Products

- 构建多图属性抽取任务,融合图文识别与跨图证据融合
- 9个大模型在产品级仅恢复49.9%属性,召回率下降15-34个百分点
- 适合工业视觉理解、多模态模型评估研究者使用
阀门、断路器等工业产品由密集的技术规范定义,这些规范分布在规格表、铭牌和工程图等异构图像中,但多模态大模型能否可靠提取仍不明确。为此,我们提出IndustryBench-MIPU,首个大规模多图工业产品理解基准,聚焦结构化属性提取——从产品图像中恢复属性-值对。该任务综合考察规格表与铭牌的文本识别、工程图的视觉推理、工业术语解析及跨图证据整合能力。基准包含4,559个产品、27,652张图像,共103,703个标注,覆盖18类工业产品,通过多模型共识与三级质量保障构建。在单图与产品级多图设置下评估9个MLLMs,结果显示:模型精度高(86–94%),但最佳模型仅恢复49.9%的产品级属性;从单图到多图提取,召回率下降15–34个百分点。多图完整性而非单图准确率是核心瓶颈。数据集与代码已公开。
原文摘要 · Abstract (English)
Industrial products such as valves and circuit breakers are defined by dense technical specifications that govern procurement, compatibility, and safety across supply chains. These specifications are scattered across multiple heterogeneous product images, including specification tables, nameplates, and technical drawings, yet whether Multimodal Large Language Models (MLLMs) can reliably recover them remains underexplored. To fill this gap, we introduce IndustryBench-MIPU, the first large-scale benchmark for multi-image industrial product understanding, built around structured attribute extraction -- recovering property-value pairs from product images. This task jointly probes text recognition on specification tables and nameplates, visual reasoning over technical drawings, domain knowledge to decode industrial terminology, and cross-image evidence integration to assemble scattered specifications. Concretely, the benchmark comprises 4,559 products across 27,652 images with 103,703 annotations spanning 18 industrial categories, constructed through multi-model consensus and three-tier quality assurance. Evaluating nine MLLMs under both single-image and product-level multi-image settings reveals a stark completeness gap: models achieve high precision (86--94%) but the best recovers only 49.9% of product-level attributes; moving from single-image to multi-image extraction costs 15--34 percentage points of recall. Multi-image completeness, not single-image accuracy, is the core bottleneck. Dataset and code are publicly available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。