系统梳理视觉识别可解释性,构建人类中心的多维分类框架。
A Survey on Interpretability in Visual Recognition
- 从意图、对象、呈现方式和方法四维度建立可解释性分类体系。
- 总结评估标准与量化指标,验证不同类别在特定维度的表现。
- 聚焦多模态大模型可解释性,适合关注AI透明化的研究者参考。
视觉识别模型在各类任务中取得了前所未有的成功。随着其在自动驾驶、医疗诊断等安全关键领域的部署需求增长,可解释人工智能(XAI)的发展加速。不同于通用XAI,视觉识别领域的XAI处于视觉与语言两大人类基础模态的交汇处,是多模态智能的核心。本文从以人为本的视角出发,建立涵盖意图、对象、呈现方式和方法的多维分类体系,系统综述了视觉识别中的可解释性研究。除分类外,还总结了关键评估标准与度量指标,对不同类别进行定性评估,并在特定维度上展示定量基准。进一步探讨了多模态大语言模型的可解释性及其实际应用,识别出新兴趋势与机遇。通过整合这些多元视角,本综述为未来视觉识别可解释性研究提供了深刻洞见与路线图。
原文摘要 · Abstract (English)
Visual recognition models have achieved unprecedented success in various tasks. While researchers aim to understand the underlying mechanisms of these models, the growing demand for deployment in safety-critical areas like autonomous driving and medical diagnostics has accelerated the development of eXplainable AI (XAI). Distinct from generic XAI, visual recognition XAI is positioned at the intersection of vision and language, which represent the two most fundamental human modalities and form the cornerstones of multimodal intelligence. This paper provides a systematic survey of XAI in visual recognition by establishing a multi-dimensional taxonomy from a human-centered perspective based on intent, object, presentation, and methodology. Beyond categorization, we summarize critical evaluation desiderata and metrics, conducting an extensive qualitative assessment across different categories and demonstrating quantitative benchmarks within specific dimensions. Furthermore, we explore the interpretability of Multimodal Large Language Models and practical applications, identifying emerging trends and opportunities. By synthesizing these diverse perspectives, this survey provides an insightful roadmap to inspire future research on the interpretability of visual recognition models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。