提出QUBA评分体系,多维度评估图像分类模型质量。
Beyond Accuracy: What Matters in Designing Well-Behaved Image Classification Models?
- 系统分析326个模型在9个质量维度的表现
- 自监督预训练显著提升多数质量指标
- 数据量是影响模型质量的关键因素
深度学习已成为计算机视觉的核心,深度神经网络(DNN)在预测性能上表现优异,但在鲁棒性、校准度或公平性等关键质量维度上常显不足。现有研究仅关注部分质量维度,未探索DNN整体“良好行为”的普遍性。本文通过大规模实验,同时考察图像分类任务中九个不同质量维度,分析326个主干模型在不同训练范式和架构下的表现。研究发现:(i) 视觉语言模型在ImageNet-1k分类中表现出高类别平衡性和强域变化鲁棒性;(ii) 使用自监督学习获得的权重初始化可有效提升大多数质量维度;(iii) 训练数据集规模是多数质量维度的主要驱动因素。最后,我们提出QUBA分数(Quality Understanding Beyond Accuracy),一种跨多维度的模型质量评分机制,支持根据用户需求进行个性化推荐。
原文摘要 · Abstract (English)
Deep learning has become an essential part of computer vision, with deep neural networks (DNNs) excelling in predictive performance. However, they often fall short in other critical quality dimensions, such as robustness, calibration, or fairness. While existing studies have focused on a subset of these quality dimensions, none have explored a more general form of "well-behavedness" of DNNs. With this work, we address this gap by simultaneously studying nine different quality dimensions for image classification. Through a large-scale study, we provide a bird's-eye view by analyzing 326 backbone models and how different training paradigms and model architectures affect these quality dimensions. We reveal various new insights such that (i) vision-language models exhibit high class balance on ImageNet-1k classification and strong robustness against domain changes; (ii) training models initialized with weights obtained through self-supervised learning is an effective strategy to improve most considered quality dimensions; and (iii) the training dataset size is a major driver for most of the quality dimensions. We conclude our study by introducing the QUBA score (Quality Understanding Beyond Accuracy), a novel metric that ranks models across multiple dimensions of quality, enabling tailored recommendations based on specific user needs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。