高效训练通用图像特征提取器,跨域表征能力强。
Efficient and Discriminative Image Feature Extraction for Universal Image Retrieval
- 构建多领域数据集M4D-35k,支持资源受限下的高效训练。
- 在谷歌通用图像嵌入挑战中达mMP@5 0.721,接近顶尖水平。
- 参数量少32%,可训练参数少289倍,适合轻量化部署。
当前图像检索系统普遍存在领域特定性和泛化能力不足的问题。本文提出一种计算高效的训练框架,用于构建跨领域强语义表征的通用特征提取器。为此,我们构建了一个多领域训练数据集M4D-35k,支持资源高效训练。同时,系统评估了多种前沿视觉-语义基础模型及基于边距的度量学习损失函数,筛选出最适合高效通用特征提取的组合。尽管计算资源受限,我们的方法在谷歌通用图像嵌入挑战赛中取得mMP@5 0.721的性能,位居排行榜第二,仅落后于最优方法0.7个百分点。此外,模型总参数减少32%,可训练参数减少289倍。相较于同类计算成本的方法,性能领先3.3个百分点。代码与M4D-35k标注数据集已开源。
原文摘要 · Abstract (English)
Current image retrieval systems often face domain specificity and generalization issues. This study aims to overcome these limitations by developing a computationally efficient training framework for a universal feature extractor that provides strong semantic image representations across various domains. To this end, we curated a multi-domain training dataset, called M4D-35k, which allows for resource-efficient training. Additionally, we conduct an extensive evaluation and comparison of various state-of-the-art visual-semantic foundation models and margin-based metric learning loss functions regarding their suitability for efficient universal feature extraction. Despite constrained computational resources, we achieve near state-of-the-art results on the Google Universal Image Embedding Challenge, with a mMP@5 of 0.721. This places our method at the second rank on the leaderboard, just 0.7 percentage points behind the best performing method. However, our model has 32% fewer overall parameters and 289 times fewer trainable parameters. Compared to methods with similar computational requirements, we outperform the previous state of the art by 3.3 percentage points. We release our code and M4D-35k training set annotations at https://github.com/morrisfl/UniFEx.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。