用通用视觉语言表征提升少样本计数的跨域泛化能力
Single Domain Generalization for Few-Shot Counting via Universal Representation Matching
- 引入预训练视觉语言模型的通用表征增强原型匹配
- 在跨域测试中性能超越现有方法,且不损失原域表现
- 适合需要应对真实场景变化的少样本计数任务
少样本计数通过极少标注样例估计图像中目标物体数量,但领域偏移严重制约其在未见场景下的泛化能力。当前方法大多遵循标准流程:从样例中提取物体原型,再与图像特征匹配生成相关图。我们指出,现有方法忽视了学习高度泛化的原型的重要性。为此,提出首个单领域泛化的少样本计数模型——通用表征匹配(URM)。核心贡献在于将大规模预训练视觉语言模型的通用视觉语言表征融入相关图构建过程,显著提升对领域偏移的鲁棒性,且不牺牲原域性能。URM在原域和新提出的领域泛化设置下均达到领先水平。
原文摘要 · Abstract (English)
Few-shot counting estimates the number of target objects in an image using only a few annotated exemplars. However, domain shift severely hinders existing methods to generalize to unseen scenarios. This falls into the realm of single domain generalization that remains unexplored in few-shot counting. To solve this problem, we begin by analyzing the main limitations of current methods, which typically follow a standard pipeline that extract the object prototypes from exemplars and then match them with image feature to construct the correlation map. We argue that existing methods overlook the significance of learning highly generalized prototypes. Building on this insight, we propose the first single domain generalization few-shot counting model, Universal Representation Matching, termed URM. Our primary contribution is the discovery that incorporating universal vision-language representations distilled from a large scale pretrained vision-language model into the correlation construction process substantially improves robustness to domain shifts without compromising in domain performance. As a result, URM achieves state-of-the-art performance on both in domain and the newly introduced domain generalization setting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。