arXiv:2511.15984cs.CV2025-11被引 1

统一检测与生成,提升电商场景细粒度视觉识别

UniDGF: A Unified Detection-to-Generation Framework for Hierarchical Object Visual Recognition

  • 用检测引导生成,分步预测类别与属性词元
  • 在电商数据集上显著优于传统相似度方法
  • 适合需要细粒度分类与属性理解的场景

实现视觉语义理解需要一个能同时处理目标检测、类别预测和属性识别的统一框架。然而,现有先进方法依赖全局相似性,在大规模电商场景中难以捕捉细粒度类别差异和类别特异的属性多样性。为此,我们提出一种检测引导的生成框架,用于预测层次化类别和属性词元。对每个检测到的物体,提取精细的ROI级特征,并采用基于BART的生成器,以粗到细的序列生成涵盖类别层级和属性-值对的语义词元,支持属性条件下的属性识别。在大规模自有电商数据集和开源数据集上的实验表明,该方法显著优于现有的基于相似度的流水线和多阶段分类系统,在细粒度识别和统一推理一致性方面表现更优。

原文摘要 · Abstract (English)

Achieving visual semantic understanding requires a unified framework that simultaneously handles object detection, category prediction, and attribute recognition. However, current advanced approaches rely on global similarity and struggle to capture fine-grained category distinctions and category-specific attribute diversity, especially in large-scale e-commerce scenarios. To overcome these challenges, we introduce a detection-guided generative framework that predicts hierarchical category and attribute tokens. For each detected object, we extract refined ROI-level features and employ a BART-based generator to produce semantic tokens in a coarse-to-fine sequence covering category hierarchies and property-value pairs, with support for property-conditioned attribute recognition. Experiments on both large-scale proprietary e-commerce datasets and open-source datasets demonstrate that our approach significantly outperforms existing similarity-based pipelines and multi-stage classification systems, achieving stronger fine-grained recognition and more coherent unified inference.

视觉识别生成模型电商应用细粒度分类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。