arXiv:2510.11091cs.CVcs.AI2025-10中稿 · The 12th Internati…

融合文字与几何信息,提升CAD图纸符号识别准确率

Text-Enhanced Panoptic Symbol Spotting in CAD Drawings

  • 联合建模几何与文本元素,构建统一表示
  • 引入类型感知注意力机制,显式捕捉元素间空间关系
  • 在复杂图纸上表现更鲁棒,适合工程设计场景

随着计算机辅助设计(CAD)在工程、建筑和工业设计中的广泛应用,准确解析和分析这些图纸的能力日益关键。其中,全景符号检测在实现CAD自动化和设计检索等下游应用中起着核心作用。现有方法主要依赖图纸中的几何图元,但普遍存在忽略丰富文本注释、缺乏对图元间关系的显式建模等问题,导致对图纸整体理解不全面。为此,我们提出一种融合文本信息的全景符号检测框架。该框架通过联合建模几何与文本图元构建统一表征,以预训练CNN提取的视觉特征为初始输入,采用带有类型感知注意力机制的Transformer主干网络,显式建模不同类型图元间的空间依赖关系。在真实世界数据集上的大量实验表明,所提方法在含文本注释的符号检测任务中优于现有方法,并在复杂CAD图纸上展现出更强的鲁棒性。

原文摘要 · Abstract (English)

With the widespread adoption of Computer-Aided Design(CAD) drawings in engineering, architecture, and industrial design, the ability to accurately interpret and analyze these drawings has become increasingly critical. Among various subtasks, panoptic symbol spotting plays a vital role in enabling downstream applications such as CAD automation and design retrieval. Existing methods primarily focus on geometric primitives within the CAD drawings to address this task, but they face following major problems: they usually overlook the rich textual annotations present in CAD drawings and they lack explicit modeling of relationships among primitives, resulting in incomprehensive understanding of the holistic drawings. To fill this gap, we propose a panoptic symbol spotting framework that incorporates textual annotations. The framework constructs unified representations by jointly modeling geometric and textual primitives. Then, using visual features extract by pretrained CNN as the initial representations, a Transformer-based backbone is employed, enhanced with a type-aware attention mechanism to explicitly model the different types of spatial dependencies between various primitives. Extensive experiments on the real-world dataset demonstrate that the proposed method outperforms existing approaches on symbol spotting tasks involving textual annotations, and exhibits superior robustness when applied to complex CAD drawings.

CAD分析符号检测多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。