arXiv:2602.13345cs.LGcs.IR2026-02

用多模态检索让老旧工程图纸自动建档,跨厂搜索更精准。

BLUEPRINT Rebuilding a Legacy: Multimodal Retrieval for Complex Engineering Drawings and Documents

  • 识别图纸区域,限定OCR范围,统一零件编号等标识
  • 在350个专家标注查询中,成功检索率提升10.1%,相关性指标提高18.9%
  • 适合处理海量无标签工程档案的自动化元数据构建

数十载的工程图纸与技术文档仍存于遗留档案中,因元数据不一致或缺失,检索困难且常需人工干预。我们提出Blueprint,一种面向大规模工程资料库的布局感知多模态检索系统。该系统检测标准绘图区域,实施区域限定的视觉语言模型OCR,对标识符(如DWG、零件、设施)进行归一化,并融合词法与密集检索,通过轻量级区域重排序器提升精度。部署于约77万份未标注文件上,可自动生成适用于跨厂检索的结构化元数据。在含350个专家标注查询的5千文件基准上评估,蓝图在Success@3上比最强视觉-语言基线提升10.1个百分点,nDCG@3相对提升18.9%,在视觉、文本及多模态意图下均表现更优。模拟理想区域检测与OCR的最优情况显示仍有显著提升空间。所有查询、结果、标注与代码均已公开,以支持对遗留工程档案的可复现评估。

原文摘要 · Abstract (English)

Decades of engineering drawings and technical records remain locked in legacy archives with inconsistent or missing metadata, making retrieval difficult and often manual. We present Blueprint, a layout-aware multimodal retrieval system designed for large-scale engineering repositories. Blueprint detects canonical drawing regions, applies region-restricted VLM-based OCR, normalizes identifiers (e.g., DWG, part, facility), and fuses lexical and dense retrieval with a lightweight region-level reranker. Deployed on ~770k unlabeled files, it automatically produces structured metadata suitable for cross-facility search. We evaluate Blueprint on a 5k-file benchmark with 350 expert-curated queries using pooled, graded (0/1/2) relevance judgments. Blueprint delivers a 10.1% absolute gain in Success@3 and an 18.9% relative improvement in nDCG@3 over the strongest vision-language baseline}, consistently outperforming across vision, text, and multimodal intents. Oracle ablations reveal substantial headroom under perfect region detection and OCR. We release all queries, runs, annotations, and code to facilitate reproducible evaluation on legacy engineering archives.

多模态检索工程文档元数据生成视觉语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。