用BEiT-3提升开放词汇全景分割性能
Open-Vocabulary Panoptic Segmentation Using BERT Pre-Training of Vision-Language Multiway Transformer Model
- 基于BEiT-3的视觉语言交叉注意力机制
- 在COCO-Panoptic和LVIS数据集上优于现有方法
- 适合需要泛化到未见类别的分割场景
开放词汇全景分割仍具挑战性,主要难点在于如何用有限标注数据使模型泛化至无限类别。当前主流方法依赖大规模视觉语言预训练基础模型,如CLIP。本文提出OMTSeg,利用另一大规模视觉语言预训练模型BEiT-3,通过挖掘其视觉与语言特征间的跨模态注意力,实现更优性能。实验表明,OMTSeg在COCO-Panoptic和LVIS数据集上均优于现有最先进模型。
原文摘要 · Abstract (English)
Open-vocabulary panoptic segmentation remains a challenging problem. One of the biggest difficulties lies in training models to generalize to an unlimited number of classes using limited categorized training data. Recent popular methods involve large-scale vision-language pre-trained foundation models, such as CLIP. In this paper, we propose OMTSeg for open-vocabulary segmentation using another large-scale vision-language pre-trained model called BEiT-3 and leveraging the cross-modal attention between visual and linguistic features in BEiT-3 to achieve better performance. Experiments result demonstrates that OMTSeg performs favorably against state-of-the-art models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。