通过上下文关系提升服装物品检测准确率
Holi-DETR: Holistic Fashion Item Detection Leveraging Contextual Information
- 融合三类上下文信息进行整体检测
- 相较基础DETR提升3.6个百分点AP
- 适合时尚图像分析与推荐系统
服装物品检测因品类外观差异大、子类相似度高而具挑战性。为此,我们提出新型全局检测变换器Holi-DETR,通过利用上下文信息在穿搭图像中整体检测服装物品。服装常以特定风格组合,具有语义关联。不同于传统独立检测每个物品的方法,Holi-DETR结合三种上下文信息:(1) 服装间的共现关系,(2) 基于物品间空间排列的相对位置与尺寸,(3) 物品与人体关键点的空间关系。为此,我们设计新架构,将这三类异构上下文信息融入检测变换器(DETR)及其后续模型。实验表明,该方法在平均精度(AP)上比原生DETR提升3.6个百分点,比近期提出的Co-DETR提升1.1个百分点。
原文摘要 · Abstract (English)
Fashion item detection is challenging due to the ambiguities introduced by the highly diverse appearances of fashion items and the similarities among item subcategories. To address this challenge, we propose a novel Holistic Detection Transformer (Holi-DETR) that detects fashion items in outfit images holistically, by leveraging contextual information. Fashion items often have meaningful relationships as they are combined to create specific styles. Unlike conventional detectors that detect each item independently, Holi-DETR detects multiple items while reducing ambiguities by leveraging three distinct types of contextual information: (1) the co-occurrence relationship between fashion items, (2) the relative position and size based on inter-item spatial arrangements, and (3) the spatial relationships between items and human body key-points. %Holi-DETR explicitly incorporates three types of contextual information: (1) the co-occurrence probability between fashion items, (2) the relative position and size based on inter-item spatial arrangements, and (3) the spatial relationships between items and human body key-points. To this end, we propose a novel architecture that integrates these three types of heterogeneous contextual information into the Detection Transformer (DETR) and its subsequent models. In experiments, the proposed methods improved the performance of the vanilla DETR and the more recently developed Co-DETR by 3.6 percent points (pp) and 1.1 pp, respectively, in terms of average precision (AP).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。