用语言引导生成模型提升工业轮廓检测精度
Generative AI for Industrial Contour Detection: A Language-Guided Vision System
- 通过条件GAN生成轮廓,再用视觉语言模型精修
- 在自研数据集上显著提升边缘连续性和几何对齐度
- 适合需要高精度轮廓检测的制造业场景
工业视觉系统常受噪声、材料差异和成像条件不可控影响,导致传统边缘检测器和手工流程效果受限。本文提出一种语言引导的生成式视觉系统,用于制造中残余轮廓检测,旨在实现接近CAD级别的精度。系统分三阶段:数据采集与预处理、基于条件GAN的轮廓生成,以及通过视觉-语言建模进行多模态轮廓精修,其中标准化提示通过人机协同过程设计,并以图像-文本引导合成方式应用。在自有FabTrack数据集上,该系统提升了轮廓保真度,增强边缘连续性与几何对齐,减少人工追踪工作量。精修阶段对比了多个视觉-语言模型,包括Google Gemini 2.0 Flash、OpenAI GPT-image-1(集成于VLM引导工作流)及开源基线。在标准条件下,GPT-image-1在结构准确性和感知质量上均持续优于Gemini 2.0 Flash。结果表明,VLM引导的生成式工作流有望突破传统工业视觉系统的局限。
原文摘要 · Abstract (English)
Industrial computer vision systems often struggle with noise, material variability, and uncontrolled imaging conditions, limiting the effectiveness of classical edge detectors and handcrafted pipelines. In this work, we present a language-guided generative vision system for remnant contour detection in manufacturing, designed to achieve CAD-level precision. The system is organized into three stages: data acquisition and preprocessing, contour generation using a conditional GAN, and multimodal contour refinement through vision-language modeling, where standardized prompts are crafted in a human-in-the-loop process and applied through image-text guided synthesis. On proprietary FabTrack datasets, the proposed system improved contour fidelity, enhancing edge continuity and geometric alignment while reducing manual tracing. For the refinement stage, we benchmarked several vision-language models, including Google's Gemini 2.0 Flash, OpenAI's GPT-image-1 integrated within a VLM-guided workflow, and open-source baselines. Under standardized conditions, GPT-image-1 consistently outperformed Gemini 2.0 Flash in both structural accuracy and perceptual quality. These findings demonstrate the promise of VLM-guided generative workflows for advancing industrial computer vision beyond the limitations of classical pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。