用视觉理解替代文字识别,自动检查工程图纸合规性
PlanSightRAG: A Visual-First Multimodal RAG for Automating Question Answering and Compliance Checking for Civil Standard Plans

- 直接分析图纸图像,融合多向量检索与智能推理流程
- 在五个州交通部门数据上达91.47%召回率,合成图验证准确率100%
- 无需人工规则,可从规范文档中自动提取数值限制,适合工程自动化场景
传统基础设施合规检查依赖工程师手动阅读2D图纸,而基于OCR的自动化方法会丢失关键的几何与布局信息。我们提出视觉优先的多模态检索增强生成框架PlanSightRAG,直接对图纸图像进行索引与推理,集成ColNomic-3B多向量检索、代理式规划-检索-审计-合成器,以及MaxSim热力图作为证据链。构建了包含4,056对样本的基准数据集,来自五个州交通部门的标准图纸(共1,898页)。PlanSightRAG在零样本检索中达到91.47% Recall@5,于密歇根州交通局数据集上实现91.40%。在参数化生成的合规图纸测试中,基于Qwen2.5-VL-72B的流水线仅在预设规则阈值下达100%判断准确率,而无视觉模型的OCR基线已达76.4%。最终演示了无需人工规则即可从规范语料库中自动提取数值限制的自主视觉规则对齐能力。
原文摘要 · Abstract (English)
Civil infrastructure compliance checking has long relied on engineers manually reading legacy 2D plans; however, OCR-based automation strips away the geometry and layout essential for interpreting these plans. We present a Visual-First Multimodal Retrieval-Augmented Generation (RAG) framework called PlanSightRAG. It indexes and reasons directly over plan imagery, integrates a ColNomic-3B multi-vector retrieval, an agentic Planner-Retriever-Auditor-Synthesizer, and MaxSim heatmaps as an evidence trail. We introduce a 4,056-pair benchmark from five state Departments of Transportation (DOT) standard plans (1,898 pages). PlanSightRAG achieves 91.47% Recall@5 on zero-shot retrieval, while on a held-out Michigan DOT corpus, it achieves 91.40%. On synthetic, parametrically-generated compliance drawings, our Qwen2.5-VL-72B pipeline reaches 100% verdict accuracy only when supplied a pre-resolved rule threshold, a controlled ceiling that a non-VLM OCR baseline already reaches at 76.4%. Finally, we demonstrate autonomous visual rule-grounding by extracting numeric limits directly from a specification corpus without any human-supplied rules.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。