无需训练,通过部件增强提升细粒度图像描述精度
TROPE: TRaining-Free Object-Part Enhancement for Seamlessly Improving Fine-Grained Zero-Shot Image Captioning
- 利用检测框与NLP技术补充图像描述中的物体部件细节
- 在多个细粒度数据集上显著提升零样本图像描述效果
- 兼容现有方法,适合需要精准描述的视觉任务
零样本推理是如CLIP等大模型的新兴能力。尽管已有研究提升主流数据集(如MSCOCO、Flickr8k)上的零样本图像描述表现,但在细粒度数据集(如CUB、FLO、UCM-Captions、Sydney-Captions)上仍不足。这些数据集要求区分视觉和语义相似类别,关注物体部件及其属性。为此,我们提出无需训练的物体部件增强方法TROPE。TROPE通过物体检测建议框与自然语言处理技术,在基础描述中添加额外部件细节,不修改原描述,可无缝集成至其他生成方法,提供更高灵活性。实验表明,TROPE在所有测试的零样本图像描述方法中均持续提升性能,并在细粒度图像描述数据集上达到最新水平。
原文摘要 · Abstract (English)
Zero-shot inference, where pre-trained models perform tasks without specific training data, is an exciting emergent ability of large models like CLIP. Although there has been considerable exploration into enhancing zero-shot abilities in image captioning (IC) for popular datasets such as MSCOCO and Flickr8k, these approaches fall short with fine-grained datasets like CUB, FLO, UCM-Captions, and Sydney-Captions. These datasets require captions to discern between visually and semantically similar classes, focusing on detailed object parts and their attributes. To overcome this challenge, we introduce TRaining-Free Object-Part Enhancement (TROPE). TROPE enriches a base caption with additional object-part details using object detector proposals and Natural Language Processing techniques. It complements rather than alters the base caption, allowing seamless integration with other captioning methods and offering users enhanced flexibility. Our evaluations show that TROPE consistently boosts performance across all tested zero-shot IC approaches and achieves state-of-the-art results on fine-grained IC datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。