arXiv:2508.20227cs.CVcs.AI2025-08

用视觉语言模型自动分析图像模型的推理逻辑,发现缺陷并理解整体行为。

A Novel Framework for Automated Explain Vision Model Using Vision-Language Models

  • 构建端到端流程,结合视觉语言模型在样本与数据集层面解释视觉模型
  • 可识别模型失败案例并揭示其在大规模数据上的行为模式
  • 适合想提升模型透明度与可靠性的开发者和研究者

许多视觉模型的发展集中于提升准确率、IoU 和 mAP 等指标,却较少关注可解释性,因传统 xAI 方法难以对训练好的模型提供有意义的解释。尽管现有 xAI 技术多针对单个样本进行解释,但对模型在大规模数据上展现的通用行为分析仍不充分。而理解模型在一般图像上的表现对于避免偏见判断、识别模型趋势与模式至关重要。本文借助视觉语言模型(Vision-Language Models),提出一个可在样本与数据集双层面解释视觉模型的新框架。该框架能以极低投入发现模型失败案例,并深入洞察模型行为,实现视觉模型开发与可解释性分析的融合,推动图像分析技术进步。

原文摘要 · Abstract (English)

The development of many vision models mainly focuses on improving their performance using metrics such as accuracy, IoU, and mAP, with less attention to explainability due to the complexity of applying xAI methods to provide a meaningful explanation of trained models. Although many existing xAI methods aim to explain vision models sample-by-sample, methods explaining the general behavior of vision models, which can only be captured after running on a large dataset, are still underexplored. Furthermore, understanding the behavior of vision models on general images can be very important to prevent biased judgments and help identify the model's trends and patterns. With the application of Vision-Language Models, this paper proposes a pipeline to explain vision models at both the sample and dataset levels. The proposed pipeline can be used to discover failure cases and gain insights into vision models with minimal effort, thereby integrating vision model development with xAI analysis to advance image analysis.

可解释性视觉语言模型模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。