arXiv:2509.25811cs.CVcs.LG2025-09

让AI看图识标,不用记住上万品牌也能准确识别新品牌。

Logo-VGR: Visual Grounded Reasoning for Open-world Logo Recognition

  • 把识标变成图片与候选标志的对比任务,避免直接记品牌名。
  • 在未知品牌测试中比基线高出近10分,泛化能力更强。
  • 适合做产品审核、品牌识别等实际场景的AI系统开发者。

多模态大模型虽在通用基准上取得进展,但在智能产品审核等专业场景应用仍不足。为此,我们提出一个开放世界商标识别基准,核心挑战在于产品审核。传统方法需记忆数万品牌特征,在真实场景不现实。本文提出的Logo-VGR方法仅需少量品牌标注即可实现大规模品牌识别。我们将商标识别重构为对比任务:模型需匹配产品图像与候选商标,而非直接生成品牌标签。我们发现现有模型易因记忆品牌分布而过拟合,导致对未见品牌表现差。为此,Logo-VGR引入新范式:通过领域知识注入(Logo Perception Grounding)和标识引导的视觉定位推理(Logo-Guided Visual Grounded Reasoning),增强模型的多模态推理能力。实验表明,该方法在跨域设置下优于强基线近10个百分点,显著提升泛化性能。

原文摘要 · Abstract (English)

Recent advances in multimodal large language models (MLLMs) have been primarily evaluated on general-purpose benchmarks, while their applications in domain-specific scenarios, such as intelligent product moderation, remain underexplored. To address this gap, we introduce an open-world logo recognition benchmark, a core challenge in product moderation. Unlike traditional logo recognition methods that rely on memorizing representations of tens of thousands of brands-an impractical approach in real-world settings-our proposed method, Logo-VGR, enables generalization to large-scale brand recognition with supervision from only a small subset of brands. Specifically, we reformulate logo recognition as a comparison-based task, requiring the model to match product images with candidate logos rather than directly generating brand labels. We further observe that existing models tend to overfit by memorizing brand distributions instead of learning robust multimodal reasoning, which results in poor performance on unseen brands. To overcome this limitation, Logo-VGR introduces a new paradigm of domain-specific multimodal reasoning: Logo Perception Grounding injects domain knowledge, and Logo-Guided Visual Grounded Reasoning enhances the model's reasoning capability. Experimental results show that Logo-VGR outperforms strong baselines by nearly 10 points in OOD settings, demonstrating superior generalization.

商标识别多模态开放世界视觉推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。