用真实拓扑生成合成图纸,让模型在无真实数据下也能精准识别工艺图管线。
SynthPID: P&ID digitization from Topology-Preserving Synthetic Data

- 基于真实图纸拓扑生成合成工艺图,保持真实布局结构。
- 仅用合成数据训练,边缘检测准确率达63.8%,接近真实数据模型。
- 揭示合成数据质量比数量更重要,400张后性能趋于饱和。
自动化将管道与仪表图(P&IDs)转化为结构化流程图可显著提升工厂运营价值,但受限于数据稀缺:工程图纸属专有信息,公开基准仅含12张标注图像。以往的合成数据增强方法因模板生成符号随机分布,导致图结构与真实工厂差异大,仅达约33%的边检测准确率。本文提出SynthPID,构建包含665张合成P&ID的语料库,其管道拓扑直接源自真实图纸。结合专为高分辨率图设计的patch-based Relationformer,仅用合成数据训练的模型在PID2Graph OPEN100上达到63.8±3.1%的边mAP,距离真实数据基线仅差8个百分点。控制实验表明,性能提升源于生成质量而非模型选择。规模研究显示,超过约400张合成图后收益趋于平缓,提示种子多样性是主要瓶颈。
原文摘要 · Abstract (English)
Automating the digitization of Piping and Instrumentation Diagrams (P&IDs) into structured process graphs would unlock significant value in plant operations, yet progress is bottlenecked by a fundamental data problem: engineering drawings are proprietary, and the entire community shares a single public benchmark of just 12 annotated images. Prior attempts at synthetic augmentation have fallen short because template-based generators scatter symbols at random, producing graphs that bear little resemblance to real process plants and, accordingly, yield only approximately 33% edge detection accuracy under synth-only training. We argue the failure is structural rather than visual and address it by introducing SynthPID, a corpus of 665 synthetic P&IDs whose pipe topology is seeded directly from real drawings. Paired with a patch-based Relationformer adapted for high-resolution diagrams, a model trained on SynthPID alone achieves 63.8 +/- 3.1% edge mAP on PID2Graph OPEN100 without seeing a single real P&ID during training, closing within 8 pp of the real-data oracle. These gains hold up under a controlled comparison against the template-based regime, confirming that generation quality drives performance rather than model choice. A scaling study reveals that gains flatten beyond roughly 400 synthetic images, pointing to seed diversity as the binding constraint.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。