arXiv:2410.16028cs.CV2024-10

无需文本描述,用少量图像即可快速实现目标检测。

Few-shot target-driven instance detection based on open-vocabulary object detection models

  • 利用视觉-文本共享空间,通过提示实现少样本检测。
  • 模型越大、样本越多、加图像增强,检测效果越好。
  • 适合资源有限但需快速适配新目标的场景。

当前大型开放视觉模型可用于零样本或少样本物体识别,但基于梯度的微调成本较高。而开放词汇物体检测模型将视觉与文本概念映射到同一潜在空间,可通过提示实现零样本检测,计算开销小。本文提出一种轻量级方法,将此类模型转化为无需文本描述的一次或少样本识别模型。在TEgO数据集上以YOLO-World为基础模型进行实验,结果表明:模型规模越大、样本数量越多、加入图像增强,性能越优。

原文摘要 · Abstract (English)

Current large open vision models could be useful for one and few-shot object recognition. Nevertheless, gradient-based re-training solutions are costly. On the other hand, open-vocabulary object detection models bring closer visual and textual concepts in the same latent space, allowing zero-shot detection via prompting at small computational cost. We propose a lightweight method to turn the latter into a one-shot or few-shot object recognition models without requiring textual descriptions. Our experiments on the TEgO dataset using the YOLO-World model as a base show that performance increases with the model size, the number of examples and the use of image augmentation.

少样本检测开放词汇视觉-语言

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。