arXiv:2510.00603cs.CV2025-10被引 1

用大模型自动标注结构缺陷,省去人工,准确率超98%

LVLMs as inspectors: an agentic framework for category-level structural defect annotation

  • 用大模型+自问自答机制,自动识别缺陷图像
  • 在平衡数据下准确率达85%-98%,不平衡数据也达80%-92%
  • 适合需要高质量缺陷数据的工程安全与模型训练场景

自动化结构缺陷标注对保障基础设施安全至关重要,同时可降低人工标注的成本与低效问题。本文提出一种基于智能体的标注框架(ADPT),融合大视觉语言模型(LVLMs)、语义模式匹配模块及迭代自问自答优化机制。通过优化领域特定提示词与递归验证流程,ADPT能将原始视觉数据转化为高质量、语义化的缺陷数据集,无需任何人工监督。实验表明,在类平衡设置下,该方法在区分缺陷与非缺陷图像上达到最高98%准确率;在四类缺陷标注中,准确率为85%-98%;在类不平衡数据集上仍保持80%-92%准确率。该框架为构建高保真数据集提供了可扩展、低成本的解决方案,有力支持下游任务如迁移学习与领域自适应在结构损伤评估中的应用。

原文摘要 · Abstract (English)

Automated structural defect annotation is essential for ensuring infrastructure safety while minimizing the high costs and inefficiencies of manual labeling. A novel agentic annotation framework, Agent-based Defect Pattern Tagger (ADPT), is introduced that integrates Large Vision-Language Models (LVLMs) with a semantic pattern matching module and an iterative self-questioning refinement mechanism. By leveraging optimized domain-specific prompting and a recursive verification process, ADPT transforms raw visual data into high-quality, semantically labeled defect datasets without any manual supervision. Experimental results demonstrate that ADPT achieves up to 98% accuracy in distinguishing defective from non-defective images, and 85%-98% annotation accuracy across four defect categories under class-balanced settings, with 80%-92% accuracy on class-imbalanced datasets. The framework offers a scalable and cost-effective solution for high-fidelity dataset construction, providing strong support for downstream tasks such as transfer learning and domain adaptation in structural damage assessment.

缺陷检测大模型应用自动标注智能巡检

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。