用视觉语言模型高效选图,提升深度学习标注效率
A Highly Efficient Diversity-based Input Selection for DNN Improvement Using VLMs
- 基于视觉语言模型设计新多样性度量CBD,计算快且效果好
- 在多个数据集上优于现有方法,选图效率接近简单不确定度法
- 适合大规模重复标注场景,尤其适合资源有限的研究者
通过微调提升深度神经网络性能需标注新收集的输入,但这一过程成本高、耗时长。为缓解此问题,近年出现输入选择方法,旨在选取少量高信息量样本进行标注。多样性选择是其中最有效的方法之一,但通常计算开销大,难以扩展到大规模输入集。为此,本文提出概念多样性(CBD),一种基于视觉语言模型的新颖高效图像输入多样性度量。实验表明,CBD与已有几何多样性(GD)高度相关,但计算时间仅为后者的极小部分。在此基础上,我们提出结合CBD与边际(Margin)不确定度的混合选择方法。在多种DNN模型、输入集、选择预算及六种先进基线上的全面评估显示,基于CBD的选择方法始终优于所有基线,能更有效地提升模型性能。此外,该方法保持极高效率,在如ImageNet等大规模数据集上,选择时间接近简单不确定性方法(如Margin)。结果验证了其有效性、计算优势及在重复、大规模输入选择场景中的可扩展性。
原文摘要 · Abstract (English)
Maintaining or improving the performance of Deep Neural Networks (DNNs) through fine-tuning requires labeling newly collected inputs, a process that is often costly and time-consuming. To alleviate this problem, input selection approaches have been developed in recent years to identify small, yet highly informative subsets for labeling. Diversity-based selection is one of the most effective approaches for this purpose. However, they are often computationally intensive and lack scalability for large input sets, limiting their practical applicability. To address this challenge, we introduce Concept-Based Diversity (CBD), a novel and highly efficient diversity metric for image inputs that leverages Vision-Language Models (VLMs). Our results show that CBD exhibits a strong correlation with Geometric Diversity (GD), an established diversity metric, while requiring only a fraction of its computation time. Building on this finding, we propose a hybrid input selection approach that combines CBD with Margin, a simple uncertainty metric. We conduct a comprehensive evaluation across a diverse set of DNN models, input sets, selection budgets, and six most effective state-of-the-art selection baselines. The results demonstrate that the CBD-based selection consistently outperforms all baselines at guiding input selection to improve the DNN model. Furthermore, the CBD-based selection approach remains highly efficient, requiring selection times close to those of simple uncertainty-based methods such as Margin, even on larger input sets like ImageNet. These results confirm not only the effectiveness and computational advantage of the CBD-based approach, particularly compared to hybrid baselines, but also its scalability in repetitive and extensive input selection scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。