用不确定性建模提升图像检索的可靠性与精度
Evidential Transformers for Improved Image Retrieval
- 引入证据推理机制,以不确定度驱动图像检索
- 在SOP和CUB-200-2011上超越现有最佳结果
- 适合需要高可靠性的图像匹配场景
我们提出证据化变压器(Evidential Transformer),一种基于不确定性的变压器模型,用于改进内容驱动的图像检索(CBIR)。本文通过将概率方法融入图像检索,实现了更稳健、可靠的性能。证据分类显著优于传统多分类训练基准,在深度度量学习中表现更优。此外,借助全局上下文视觉变换器(GC ViT)架构,我们在多个数据集上提升了当前最优检索性能。实验结果一致表明该方法在斯坦福在线产品(SOP)和CUB-200-2011数据集的所有测试设置中均表现出可靠性,树立了新的基准。
原文摘要 · Abstract (English)
We introduce the Evidential Transformer, an uncertainty-driven transformer model for improved and robust image retrieval. In this paper, we make several contributions to content-based image retrieval (CBIR). We incorporate probabilistic methods into image retrieval, achieving robust and reliable results, with evidential classification surpassing traditional training based on multiclass classification as a baseline for deep metric learning. Furthermore, we improve the state-of-the-art retrieval results on several datasets by leveraging the Global Context Vision Transformer (GC ViT) architecture. Our experimental results consistently demonstrate the reliability of our approach, setting a new benchmark in CBIR in all test settings on the Stanford Online Products (SOP) and CUB-200-2011 datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。