arXiv:2508.01921cs.CV2025-08ICCV

统一视觉语言模型在工业检测中表现不佳,缺乏可靠性。

InspectVLM: Unified in Theory, Unreliable in Practice

  • 用Florence-2构建统一模型,基于新数据集InspectMM训练
  • 分类和关键点任务表现尚可,但检测精度远低于传统模型
  • 对提示变化敏感,常输出记忆内容,不适合作业场景

统一视觉语言模型(VLM)有望通过单一语言接口整合分类、检测和关键点定位等视觉任务,简化工业检测流程。本文提出InspectVLM,基于Florence-2,在我们自建的大规模多模态多任务检测数据集InspectMM上训练。尽管在图像分类和结构化关键点任务上表现良好,但在核心检测指标上仍不及传统ResNet模型。模型在低提示多样性下表现脆弱,细粒度检测产生退化输出,且频繁忽略视觉输入,直接返回记忆的语言响应。结果表明,当前语言驱动的统一架构虽概念优雅,但缺乏足够的视觉感知与鲁棒性,难以应用于高精度要求的工业检测场景。

原文摘要 · Abstract (English)

Unified vision-language models (VLMs) promise to streamline computer vision pipelines by reframing multiple visual tasks such as classification, detection, and keypoint localization within a single language-driven interface. This architecture is particularly appealing in industrial inspection, where managing disjoint task-specific models introduces complexity, inefficiency, and maintenance overhead. In this paper, we critically evaluate the viability of this unified paradigm using InspectVLM, a Florence-2-based VLM trained on InspectMM, our new large-scale multimodal, multitask inspection dataset. While InspectVLM performs competitively on image-level classification and structured keypoint tasks, we find that it fails to match traditional ResNet-based models in core inspection metrics. Notably, the model exhibits brittle behavior under low prompt variability, produces degenerate outputs for fine-grained object detection, and frequently defaults to memorized language responses regardless of visual input. Our findings suggest that while language-driven unification offers conceptual elegance, current VLMs lack the visual grounding and robustness necessary for deployment in precision critical industrial inspections.

视觉语言模型工业检测模型鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。