评测工业级图像产品检索模型,揭示现有效果与局限。
Visual Product Search Benchmark
- 统一协议下测试开源、专有及领域模型的实例级检索能力。
- 在制造、汽车等真实工业数据集上验证,准确率差异显著。
- 适合关注工业视觉识别落地的工程师与研究者参考。
从图像中可靠识别产品是工业与商业应用中的关键需求,尤其在维护、采购和运营流程中,错误匹配可能导致高昂的后续损失。系统核心是视觉搜索组件,需在多样成像条件下从大规模且持续更新的目录中检索并排序精确的物体实例。本报告构建了一个面向工业应用的现代视觉嵌入模型实例级图像检索结构化基准。评估了开源基础嵌入模型、专有跨模态嵌入系统及领域专用视觉模型,采用统一的图像到图像检索协议。基准包含来自制造业、汽车、家装、零售等生产部署的工业数据集,以及成熟的公开基准。评估不使用后处理,仅考察模型本身检索能力。结果揭示了当前基础与统一嵌入模型在细粒度实例检索任务上的迁移性能,及其与专为工业应用训练模型的对比。通过强调现实约束、异构成像条件与精确实例匹配要求,该基准旨在为从业者与研究人员提供当前视觉嵌入方法在生产级产品识别系统中的优劣洞察。配套交互式网站(https://benchmark.nyris.io)提供结果、评估细节与可视化支持。
原文摘要 · Abstract (English)
Reliable product identification from images is a critical requirement in industrial and commercial applications, particularly in maintenance, procurement, and operational workflows where incorrect matches can lead to costly downstream failures. At the core of such systems lies the visual search component, which must retrieve and rank the exact object instance from large and continuously evolving catalogs under diverse imaging conditions. This report presents a structured benchmark of modern visual embedding models for instance-level image retrieval, with a focus on industrial applications. A curated set of open-source foundation embedding models, proprietary multi-modal embedding systems, and domain-specific vision-only models are evaluated under a unified image-to-image retrieval protocol. The benchmark includes curated datasets, which includes industrial datasets derived from production deployments in Manufacturing, Automotive, DIY, and Retail, as well as established public benchmarks. Evaluation is conducted without post-processing, isolating the retrieval capability of each model. The results provide insight into how well contemporary foundation and unified embedding models transfer to fine-grained instance retrieval tasks, and how they compare to models explicitly trained for industrial applications. By emphasizing realistic constraints, heterogeneous image conditions, and exact instance matching requirements, this benchmark aims to inform both practitioners and researchers about the strengths and limitations of current visual embedding approaches in production-level product identification systems. An interactive companion website presenting the benchmark results, evaluation details, and additional visualizations is available at https://benchmark.nyris.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。