arXiv:2607.18997cs.CVcs.CE2026-07

评测五种深度学习模型在工程图纸布局检测中的表现,发现专用模型更优。

Benchmarking Deep Learning Approaches for AEC Engineering Drawing Layout Detection and Information Extraction

  • 构建专用AEC图纸布局数据集,对比五种深度学习架构
  • RF-DETR在mAP50上达0.949,Qwen3-VL的F1-score为0.911
  • 通用文档模型因领域干扰性能下降,适合自动化图纸信息提取

建筑、工程与施工(AEC)图纸的信息提取仍受限于人工效率低下,而作为组织图形与文本层级的关键中间步骤——布局检测尚未得到充分研究。现有通用文档布局模型以文本为中心优化,未在工程图纸上验证有效性。本研究构建了定制化的AEC专用布局数据集,并对五种深度学习架构进行基准测试。RF-DETR在mAP_{50}上达到0.949,创下最优表现;视觉-语言模型Qwen3-VL的F1-score达0.911,位居领先。相反,基于通用文档数据集预训练的模型因存在‘领域干扰’,导致性能下降。该研究为AEC领域的自动化信息提取奠定了坚实技术基础。

原文摘要 · Abstract (English)

Information Extraction (IE) from Architecture, Engineering, and Construction (AEC) drawings remains hindered by manual inefficiency, while Layout Detection, a vital 'middleware' organizing graphical and textual hierarchies, is underexplored. General document layout models, optimized for text-centric content, lack validation on engineering drawings. This study constructs a custom AEC-specific layouts dataset and benchmarks five deep learning architectures. RF-DETR achieves state-of-the-art performance with an $mAP_{50}$ of 0.949, while the Vision-Language Model Qwen3-VL attains a leading F1-score of 0.911. Conversely, models pre-trained on general document datasets suffer from "domain interference", causing performance degradation. This establishes a robust technical foundation for automated IE in AEC.

布局检测工程图纸深度学习信息提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。