arXiv:2502.12360cs.CVcs.AI2025-02被引 1

用大模型生成图像语义标签,发现视觉模型在人类可理解维度上的系统性弱点

Detecting Systematic Weaknesses in Vision Models along Predefined Human-Understandable Dimensions

  • 结合大模型零样本分类生成图像语义标签,构建可解释的弱化数据子集
  • 在真实和合成数据上验证,能有效识别出符合安全/领域专家定义维度的性能短板
  • 适合关注模型鲁棒性、安全评估的研究者与工程团队使用

切片发现方法(SDMs)是检测深度神经网络(DNN)系统性弱点的重要工具,能够识别出在特定数据子集中性能低下的语义一致片段。为使切片结果对实际应用有价值,其应与人类可理解且相关的维度对齐,例如由安全或领域专家定义的操作设计域(ODD)。尽管现有方法在结构化数据上表现良好,但在图像数据上因缺乏语义元数据而受限。为此,本文提出一种算法,将基础模型用于零样本图像分类以生成语义元数据,并结合组合搜索方法发现图像中的系统性弱点。相比现有方法,本方法识别的弱化切片更符合预定义的人类可理解维度。由于引入了基础模型,中间与最终结果可能存在噪声,因此我们进一步设计了处理噪声元数据影响的方法。我们在合成及真实世界数据集上验证了该算法的有效性,证明其能够恢复出具有人类可解释性的系统性弱点。此外,通过该方法,我们还发现了多个公开可用的顶尖视觉DNN在不同条件下的系统性缺陷。

原文摘要 · Abstract (English)

Slice discovery methods (SDMs) are prominent algorithms for finding systematic weaknesses in DNNs. They identify top-k semantically coherent slices/subsets of data where a DNN-under-test has low performance. For being directly useful, slices should be aligned with human-understandable and relevant dimensions, which, for example, are defined by safety and domain experts as part of the operational design domain (ODD). While SDMs can be applied effectively on structured data, their application on image data is complicated by the lack of semantic metadata. To address these issues, we present an algorithm that combines foundation models for zero-shot image classification to generate semantic metadata with methods for combinatorial search to find systematic weaknesses in images. In contrast to existing approaches, ours identifies weak slices that are in line with pre-defined human-understandable dimensions. As the algorithm includes foundation models, its intermediate and final results may not always be exact. Therefore, we include an approach to address the impact of noisy metadata. We validate our algorithm on both synthetic and real-world datasets, demonstrating its ability to recover human-understandable systematic weaknesses. Furthermore, using our approach, we identify systematic weaknesses of multiple pre-trained and publicly available state-of-the-art computer vision DNNs.

模型评测视觉模型系统性弱点零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。