融合图文信息提升CAD图纸符号识别准确率
Text-Aided Multi-Modal Panoptic Symbol Spotting for CAD Floor Plan Drawings

- 设计双模态编码器,联合建模图形元素与文本属性
- 在真实建筑图纸数据集上达到当前最佳性能
- 适合需要精准识别工程图标的工业自动化场景
计算机辅助设计(CAD)平面图包含图形元素和文本注释,二者提供互补的几何与语义线索,对智能设计理解至关重要。在各类CAD分析任务中,全景符号定位因工业数字化需求增长而日益重要。然而,现有方法多以图形元素为中心,忽视文本注释的语义价值。即便少数文本感知方法也仅浅层处理注释,未能充分建模其复杂语法与层次化语义,导致语义丢失与定位性能下降。为此,本文提出TextCAD,一种联合建模图形元素与文本注释的多模态框架。具体地,设计类型-属性关联编码器(TACE),通过联合建模注释类型与属性来显式编码其组合语义;进一步引入语义层级对齐框架,结合多级语义过滤(MSF)与图形元素下采样,自适应地在不同语义层级对齐注释语义与图形元素,实现跨模态语义注入与融合。在真实建筑图纸数据集上的实验表明,TextCAD显著提升符号定位性能,达到当前最优结果。
原文摘要 · Abstract (English)
Computer-Aided Design (CAD) floor plan drawings contain both graphical primitives and textual annotations, which provide complementary geometric and semantic cues for intelligent design understanding. Among CAD analysis tasks, panoptic symbol spotting has become increasingly important with the growing demand for industrial digitalization and deep learning-based automation. However, most existing methods remain primarily primitive-centric and underexploit textual annotations, despite their critical semantic value. Even the few text-aware approaches often treat annotations only superficially, without properly modeling complex syntax and hierarchical semantics of CAD annotations, which leads to semantic loss and suboptimal spotting performance. To address these limitations, we propose TextCAD, a multimodal framework that jointly models graphical primitives and textual annotations for panoptic symbol spotting. Specifically, we design a Type-Attribute Correlation Encoder (TACE) to explicitly encode the compositional semantics within annotations by jointly modeling their types and attributes. We further introduce a Semantic Hierarchy Alignment framework with Multi-level Semantic Filtering (MSF) and primitive downsampling, which adaptively aligns annotation semantics with graphical primitives at different semantic levels and enables accurate cross-modal semantic injection and fusion. Experiments on real-world building-design datasets show that TextCAD effectively improves symbol spotting performance and achieves state-of-the-art results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。