将人类视觉的共线性原理引入计算机视觉,提升工业缺陷检测性能。
On the Transfer of Collinearity to Computer Vision
- 基于人类视觉共线性机制设计新模型,增强对直线结构的感知能力。
- 在晶圆缺陷检测中错误率降低1.24倍(6.5%→5.26%),纳米材料识别提升3.2倍。
- 适合人工结构图像(如芯片、纳米材料),不适用于复杂自然图像如ImageNet。
共线性是人脑视觉感知中一种强化沿直线排列边缘的现象。然而,其在现实世界中的作用尚不明确,且在计算机视觉与工程应用中几乎未被探索。本文旨在将共线性原理迁移至计算机视觉,通过原型模型系统验证其在四个应用场景中的潜力:结合深度学习进行草图分析(案例Ⅰ、Ⅱ)、与显著性模型结合(案例Ⅱ),以及作为特征检测器(案例Ⅰ)。实验表明,在晶圆缺陷检测中,共线性使错误率从6.5%降至5.26%,性能提升1.24倍;在纳米材料缺陷识别中,错误率从21.65%降至6.64%,性能提升3.2倍。第三项实验处理遮挡问题,第四项在ImageNet上测试发现共线性效果有限。结果表明,共线性适用于具有明显人工直线结构的场景(如晶圆、纳米材料),而对自然图像(如ImageNet)帮助较小。研究揭示该原理在工业视觉任务中的适用性,为计算机视觉提供新工具。
原文摘要 · Abstract (English)
Collinearity is a visual perception phenomenon in the human brain that amplifies spatially aligned edges arranged along a straight line. However, it is vague for which purpose humans might have this principle in the real-world, and its utilization in computer vision and engineering applications even is a largely unexplored field. In this work, our goal is to transfer the collinearity principle to computer vision, and we explore the potential usages of this novel principle for computer vision applications. We developed a prototype model to exemplify the principle, then tested it systematically, and benchmarked it in the context of four use cases. Our cases are selected to spawn a broad range of potential applications and scenarios: sketching the combination of collinearity with deep learning (case I and II), using collinearity with saliency models (case II), and as a feature detector (case I). In the first use case, we found that collinearity is able to improve the fault detection of wafers and obtain a performance increase by a factor 1.24 via collinearity (decrease of the error rate from 6.5% to 5.26%). In the second use case, we test the defect recognition in nanotechnology materials and achieve a performance increase by 3.2x via collinearity (deep learning, error from 21.65% to 6.64%), and also explore saliency models. As third experiment, we cover occlusions; while as fourth experiment, we test ImageNet and observe that it might not be very beneficial for ImageNet. Therefore, we can assemble a list of scenarios for which collinearity is beneficial (wafers, nanotechnology, occlusions), and for what is not beneficial (ImageNet). Hence, we infer collinearity might be suitable for industry applications as it helps if the image structures of interest are man-made because they often consist of lines. Our work provides another tool for CV, hope to capture the power of human processing.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。