无需文本描述,用少量图像即可快速实现目标检测。
Few-shot target-driven instance detection based on open-vocabulary object detection models
- 利用视觉-文本共享空间,通过提示实现少样本检测。
- 模型越大、样本越多、加图像增强,检测效果越好。
- 适合资源有限但需快速适配新目标的场景。
当前大型开放视觉模型可用于零样本或少样本物体识别,但基于梯度的微调成本较高。而开放词汇物体检测模型将视觉与文本概念映射到同一潜在空间,可通过提示实现零样本检测,计算开销小。本文提出一种轻量级方法,将此类模型转化为无需文本描述的一次或少样本识别模型。在TEgO数据集上以YOLO-World为基础模型进行实验,结果表明:模型规模越大、样本数量越多、加入图像增强,性能越优。
原文摘要 · Abstract (English)
Current large open vision models could be useful for one and few-shot object recognition. Nevertheless, gradient-based re-training solutions are costly. On the other hand, open-vocabulary object detection models bring closer visual and textual concepts in the same latent space, allowing zero-shot detection via prompting at small computational cost. We propose a lightweight method to turn the latter into a one-shot or few-shot object recognition models without requiring textual descriptions. Our experiments on the TEgO dataset using the YOLO-World model as a base show that performance increases with the model size, the number of examples and the use of image augmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。