arXiv:2607.09888cs.CV2026-07

用多阶段对比学习解决商品目录与真实图片的匹配难题。

Bridging the Catalog-to-Real Gap: Scalable Product Recognition via Multi-Stage Contrastive Learning

论文配图:Bridging the Catalog-to-Real Gap: Scalable Product Recognition via Multi-Stage Contrastive Learning
图 1 · 摘自论文原文
  • 将识别任务转为跨域检索,通过对比学习对齐目录与实拍图像特征
  • 在无真实训练图情况下仍实现优秀零样本泛化性能
  • 适合大规模零售场景中快速适配新商品的系统

自动化商品识别是现代零售智能的核心;然而,准确匹配真实门店图像与企业庞大商品目录,仍是大规模应用中的主要可扩展性瓶颈。本文将该任务重新定义为基于嵌入的跨域检索问题,而非标准的闭集分类。具体而言,目标是从海量库存中为给定的真实商品图像片段检索最匹配的目录参考图。为弥合高质量展厅图与嘈杂门店图之间的严重领域差异,我们提出一种新颖的目录到真实多阶段对比学习框架(Cat2Real)。该框架通过系统利用物品级和图像级相似性,驱动有针对性的困难负样本挖掘,对视觉主干网络进行微调。大量实证评估表明,该方法可无缝扩展至未见商品与类别,在完全缺乏新商品真实训练图像的情况下,仍表现出卓越的零样本泛化能力。

原文摘要 · Abstract (English)

Automated product recognition is a cornerstone of modern retail intelligence; however, accurately matching real-world, in-store images against extensive corporate catalogs remains a major scalability bottleneck for large-scale applications. In this work, we address this challenge by reformulating the task as an embedding-based cross-domain retrieval problem rather than a standard closed-set classification task. Specifically, we define the objective as retrieving the most corresponding catalog reference image for a given real-world product query crop from an expansive inventory. To bridge the severe domain gap between pristine studio packshots and noisy in-store queries, we introduce a novel catalog-to-real multi-stage contrastive learning paradigm (Cat2Real). This framework fine-tunes a vision backbone by systematically exploiting both item-level and image-level similarities to drive targeted hard negative mining. Extensive empirical evaluations demonstrate that our paradigm scales seamlessly to unseen products and categories, yielding outstanding zero-shot generalization performance even in the complete absence of real-world training images for novel inventory.

商品识别对比学习零样本零售智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。