arXiv:2606.25277cs.ROcs.CV2026-06

用光电子神经网络实现低数据缺陷检测,大幅减少标注和数据量。

An Integrated Hardware-Software Design for Low-Data Spatial Defect Detection in Robotic Visual Inspection with Hybrid Optoelectronic Neural Networks

论文配图:An Integrated Hardware-Software Design for Low-Data Spatial Defect Detection in Robotic Visual Inspection with Hybrid Optoelectronic Neural Networks
图 1 · 摘自论文原文
  • 用数字微镜器件构建光学卷积层,直接在光域提取特征。
  • 压缩感知使数据量降90%,计算负载减60%,精度相当。
  • 用自然语言引导模型定位缺陷,免去繁琐形状标注。

为解决机器人视觉检测中数据过载和形状级标注低效问题,本文提出软硬件一体化的光电架构。采用非成像、低数据范式以降低标注依赖:首先,通过传感器闭环策略将数字微镜装置(DMD)重构为物理光学卷积层,实现光域特征提取,统一传感与处理;其次,采用基于块的压缩感知策略,将空间信息编码为低维时序信号,显著减少冗余。为避免人工缺陷形状标注,利用对比语言-图像预训练(CLIP)的通用特征,通过自然语言描述引导网络注意力聚焦于缺陷形态。同时提出定位准确率(LAA)指标量化形状级定位性能。透明材料缺陷检测实验验证系统有效性。参数分析显示测量矩阵、压缩比与块大小对精度有影响。结果表明,相比传统成像,该架构在保持同等精度下,使视觉变换器(Vision Transformer)数据量减少90%,卷积神经网络(CNN)计算负载降低60%。该低数据范式为海量数据流、高采集成本或边缘资源受限的工业自动化场景提供高效解决方案。

原文摘要 · Abstract (English)

To address data overload and inefficient shape-level annotation in robotic visual inspection, this paper proposes a hardware-software integrated optoelectronic architecture. A non-imaging, low-data paradigm is established to minimize annotation dependency. First, a sensor-in-the-loop strategy reconfigures a Digital Micromirror Device (DMD) as a physical optical convolutional layer, enabling photonic-domain feature extraction that unifies sensing hardware and processing software. To suppress data volume at the source, a block-based compressed sensing strategy encodes spatial information into low-dimensional temporal signals, drastically reducing redundancy. Subsequently, to bypass laborious manual defect shape annotation, natural language descriptions guide the network to align with highly generalizable features from Contrastive Language-Image Pre-training (CLIP), steering the attention maps of the optoelectronic neural network toward defect shapes. Furthermore, a Localization Accuracy for Attention (LAA) metric is proposed to quantify shape-level defect localization performance. Experiments on transparent material defect detection validate the system's effectiveness. Parametric analysis reveals how measurement matrices, compression ratios, and block sizes affect accuracy. Results show that, compared to traditional imaging, the proposed architecture maintains equivalent accuracy while reducing data volume by 90% for Vision Transformers and computational workload by 60% for Convolutional Neural Networks. This low-data paradigm offers an efficient solution for industrial automation scenarios involving massive data streams, high acquisition costs, or constrained edge resources.

缺陷检测光电子低数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。