工业视觉仿真到真实迁移的成败,关键看是否有可用先验信息。
Prior Availability in Industrial Visual Sim-to-Real: A Review of CAD-Guided and CAD-Unavailable Regimes

- 按先验信息有无划分三类场景:有CAD、无CAD、部分先验。
- 实验显示仅靠渲染数量无法提升迁移效果,需关注分布设计和校准细节。
- 有无CAD决定了验证方式:几何一致性或特征偏离度检测。
工业视觉仿真到真实世界的迁移常被理解为从合成图像到真实图像的转换,但实际部署中存在更广泛的证据与决策之间的不匹配。系统可能基于CAD渲染、模拟RGB-D数据、正常参考图、合成缺陷、预训练特征空间或语言提示构建,却需在不同传感器、光照、材料、夹具、标定误差、生产变异及罕见缺陷模式下运行。本文将该问题重新定义为先验可用性驱动的域差距问题,区分三种情形:有CAD时可支持渲染、标定、位姿估计、分割及测试时几何验证;无CAD时则依赖正常外观、特征分布、师生残差、合成异常假设、基础特征或视觉-语言先验;边界先验则保留部分CAD功能,如近似模型、模板、参考视图或语义对应。该框架连接了基于CAD的检测与6D位姿估计文献,以及通常独立评审的工业异常与表面检测研究。通过T-LESS/BOP、MVTec AD和VisA三个数据集实证锚点,发现渲染数量本身不足以实现有效迁移,源分布设计、检测器容量与小样本真实校准更为关键。同时表明,测试时使用CAD可形成掩码、位姿与深度一致性验证通道,而无CAD的检测则依赖校准后的正常性与特征偏差判断。因此,反对单一跨任务排行榜,主张应关注部署决策所依赖的先验基础。
原文摘要 · Abstract (English)
Industrial visual sim-to-real is often described as transferring from synthetic images to real images, but industrial deployment usually involves a broader mismatch between available evidence and required decisions. A system may be built from CAD renderings, simulated RGB-D observations, normal reference images, synthetic defects, pretrained feature spaces, or language prompts, yet deployed under different sensors, lighting, materials, fixtures, calibration, production variation, and rare defect modes. This review reframes industrial visual sim-to-real as a domain-gap problem organized by prior availability. We distinguish CAD-available settings, where explicit object geometry can support rendering, calibration, pose estimation, segmentation, and test-time geometric verification; CAD-unavailable settings, where geometry is replaced by normal-reference appearance, feature distributions, teacher-student residuals, synthetic anomaly assumptions, foundation features, or vision-language priors; and boundary-prior settings, where approximate models, templates, reference views, or semantic correspondences preserve only part of the CAD role. This framing connects CAD-based detection and 6D pose-estimation literature with industrial anomaly and surface-inspection literature that is usually reviewed separately. To make the taxonomy concrete, we use empirical anchors on T-LESS/BOP, MVTec AD, and VisA. The anchors show that CAD render count alone does not close transfer; source-distribution design, detector capacity, and small real calibration can matter more. They also show that CAD at test time creates a distinct verification channel through mask, pose, and depth consistency, whereas CAD-unavailable inspection relies on calibrated normality and feature deviation. The review therefore argues against a single cross-task leaderboard and instead asks what prior grounds the deployment decision.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。