解决细粒度零样本目标检测难题,提升相似物种识别能力
Fine-Grained Zero-Shot Object Detection
- 设计多层级语义对齐损失,强化视觉与语义空间关联
- 在1432个鸟种上实现超越现有模型的检测精度
- 构建首个细粒度零样本检测数据集,支持复杂分类任务
零样本目标检测旨在利用语义描述定位并识别已见和未见类别的物体。现有方法主要针对类别间差异明显的粗粒度检测任务,但在真实场景中,如鸟类、鱼类和花卉等细粒度识别问题更为常见。本文提出并解决新的细粒度零样本目标检测(FG-ZSD)问题,即在类别间细节差异极小的条件下进行检测。为此,我们提出MSHC方法,基于改进的两阶段检测器,并引入多层级语义感知嵌入对齐损失,增强视觉与语义空间的紧密耦合。由于现有ZSD数据集不适用于该任务,我们构建了首个FG-ZSD基准数据集FGZSD-Birds,包含148,820张图像,涵盖36个目、140个科、579个属和1432个种。在该数据集上的大量实验表明,所提方法显著优于现有ZSD模型。
原文摘要 · Abstract (English)
Zero-shot object detection (ZSD) aims to leverage semantic descriptions to localize and recognize objects of both seen and unseen classes. Existing ZSD works are mainly coarse-grained object detection, where the classes are visually quite different, thus are relatively easy to distinguish. However, in real life we often have to face fine-grained object detection scenarios, where the classes are too similar to be easily distinguished. For example, detecting different kinds of birds, fishes, and flowers. In this paper, we propose and solve a new problem called Fine-Grained Zero-Shot Object Detection (FG-ZSD for short), which aims to detect objects of different classes with minute differences in details under the ZSD paradigm. We develop an effective method called MSHC for the FG-ZSD task, which is based on an improved two-stage detector and employs a multi-level semantics-aware embedding alignment loss, ensuring tight coupling between the visual and semantic spaces. Considering that existing ZSD datasets are not suitable for the new FG-ZSD task, we build the first FG-ZSD benchmark dataset FGZSD-Birds, which contains 148,820 images falling into 36 orders, 140 families, 579 genera and 1432 species. Extensive experiments on FGZSD-Birds show that our method outperforms existing ZSD models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。