通过迭代反馈优化用户自然语言目标描述,提升开放词汇目标检测效果。
An Iterative Feedback Mechanism for Improving Natural Language Class Descriptions in Open-Vocabulary Object Detection
- 利用文本嵌入分析与对比样本嵌入融合改进描述
- 在多个公开模型上验证性能显著提升
- 适合非技术人员快速定义新目标类别
开放词汇目标检测模型的进步将使自动目标识别系统更可持续,并让非技术用户在多种应用场景中灵活复用。新类别可通过现场输入自然语言描述即时定义,无需重新训练模型。本文提出一种方法,通过分析文本嵌入并合理组合对比样本的嵌入,改进非技术人员对目标的自然语言描述。我们通过多个公开的开放词汇目标检测模型验证了该反馈机制带来的性能提升。
原文摘要 · Abstract (English)
Recent advances in open-vocabulary object detection models will enable Automatic Target Recognition systems to be sustainable and repurposed by non-technical end-users for a variety of applications or missions. New, and potentially nuanced, classes can be defined with natural language text descriptions in the field, immediately before runtime, without needing to retrain the model. We present an approach for improving non-technical users' natural language text descriptions of their desired targets of interest, using a combination of analysis techniques on the text embeddings, and proper combinations of embeddings for contrastive examples. We quantify the improvement that our feedback mechanism provides by demonstrating performance with multiple publicly-available open-vocabulary object detection models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。