用语言层次提升弱监督目标检测的泛化能力
Open-Vocabulary Object Detection via Language Hierarchy
- 引入语言层次结构扩展图像标签,缓解标签不匹配问题
- 在14个数据集上实现一致更优的泛化性能
- 适合需要跨域泛化的目标检测研究者
近期关于可泛化目标检测的研究受到越来越多关注,其利用大规模带图像级标签的数据集提供额外弱监督。然而,弱监督检测学习常面临图像与框标签不匹配的问题,即图像级标签无法提供精确物体信息。为此,本文设计了语言层次自训练(LHST),将语言层次引入弱监督检测器训练中,以学习更具泛化能力的检测器。LHST通过语言层次扩展图像级标签,并实现扩展标签与自训练之间的协同正则化:扩展标签提供更丰富监督并缓解标签不匹配,而自训练则根据预测可靠性评估和筛选扩展标签。此外,还设计了语言层次提示生成方法,将语言层次引入提示生成,有助于弥合训练与测试间的词汇差距。大量实验表明,所提技术在14个广泛研究的目标检测数据集上均取得优越且一致的泛化性能。
原文摘要 · Abstract (English)
Recent studies on generalizable object detection have attracted increasing attention with additional weak supervision from large-scale datasets with image-level labels. However, weakly-supervised detection learning often suffers from image-to-box label mismatch, i.e., image-level labels do not convey precise object information. We design Language Hierarchical Self-training (LHST) that introduces language hierarchy into weakly-supervised detector training for learning more generalizable detectors. LHST expands the image-level labels with language hierarchy and enables co-regularization between the expanded labels and self-training. Specifically, the expanded labels regularize self-training by providing richer supervision and mitigating the image-to-box label mismatch, while self-training allows assessing and selecting the expanded labels according to the predicted reliability. In addition, we design language hierarchical prompt generation that introduces language hierarchy into prompt generation which helps bridge the vocabulary gaps between training and testing. Extensive experiments show that the proposed techniques achieve superior generalization performance consistently across 14 widely studied object detection datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。