多视角融合语义与显著性,提升反光表面缺陷检测准确率
Multi-View Reflective Surface Inspection via Semantic-Saliency Cross-Verification

- 每视角独立处理,用视觉语言模型生成语义框,重建分支生成显著图
- 跨视角对齐语义与显著性,重排提案使AP50从52.6%提至62.6%
- 无需视角配准,三视角融合使产品级召回率从75.5%升至88.3%
智能手机背盖玻璃等反射表面的缺陷检测因视角依赖和镜面反射而困难,同一缺陷在不同视角下可见性差异大,且视觉证据存在空间模糊。为此,我们提出一种多视角检测框架:每个RGB图像由共享的视点专家处理,视觉语言模型(VLM)生成类相关语义框,法向量参考重建分支提供类无关显著性图。两者的空间一致性作为支持证据,用于重排语义提案,不修改其坐标也不将显著性视为真值。最终在产品层面融合跨视角证据,无需跨视图注册。在282张产线图像上,语义-显著性关联使AP50从52.6%提升至62.6%;在94种产品上,使用三视角时产品级召回率[email protected]从单视角最优的75.5%提升至88.3%。结果验证了语义-显著性交叉验证与额外光学观测的互补价值。
原文摘要 · Abstract (English)
Reflective smartphone cover glass is challenging to inspect from a single fixed viewpoint because defect visibility varies with viewing geometry and specular reflections. This gives rise to two practical challenges: defects may be weakly observable from certain viewpoints, while the available visual evidence may remain spatially ambiguous. To address these issues, we propose a multi-view inspection framework in which each RGB observation is processed by a shared per-view expert. A vision-language model (VLM) produces class-aware semantic boxes, while a normal-reference reconstruction branch provides class-agnostic saliency. Their spatial agreement is used as supporting evidence to rank semantic proposals without modifying their coordinates or treating saliency as ground truth. The resulting evidence records are combined at product level without cross-view registration. On 282 production-line images, semantic-saliency association improves $AP_{50}$ from 52.6% to 62.6% by re-ranking fixed semantic proposals. Across 94 products, cross-view evidence recall $R_{\rm prod}@0.5$ increases from 75.5% for the best single view to 88.3% using all three views. These results support the complementary roles of semantic-saliency cross-verification and additional optical observations in reflective-surface inspection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。