arXiv:2510.03376cs.CVeess.IV2025-10被引 1

用视觉语言模型自动评估工业图中目标检测质量并优化结果

Visual Language Model as a Judge for Object Detection in Industrial Diagrams

  • 用视觉语言模型分析工业图检测结果的完整性与一致性
  • 在复杂工业图上实现检测准确率提升,减少漏检与误检
  • 适合数字孪生、智能工业自动化领域研究者参考

工业图如管道仪表图(P&IDs)对工业设施的设计、运行和维护至关重要。将这些图纸数字化是构建数字孪生和实现智能工业自动化的关键一步。该过程的核心挑战在于精准的目标检测。尽管近年来目标检测算法取得显著进展,但缺乏自动评估检测结果质量的方法。本文提出一种利用视觉语言模型(VLMs)评估检测结果并指导优化的框架。该方法利用VLM的多模态能力识别遗漏或不一致的检测项,实现自动化质量评估,并提升复杂工业图上的整体检测性能。

原文摘要 · Abstract (English)

Industrial diagrams such as piping and instrumentation diagrams (P&IDs) are essential for the design, operation, and maintenance of industrial plants. Converting these diagrams into digital form is an important step toward building digital twins and enabling intelligent industrial automation. A central challenge in this digitalization process is accurate object detection. Although recent advances have significantly improved object detection algorithms, there remains a lack of methods to automatically evaluate the quality of their outputs. This paper addresses this gap by introducing a framework that employs Visual Language Models (VLMs) to assess object detection results and guide their refinement. The approach exploits the multimodal capabilities of VLMs to identify missing or inconsistent detections, thereby enabling automated quality assessment and improving overall detection performance on complex industrial diagrams.

目标检测视觉语言模型工业图

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。