让视觉模型解释更互动,提升理解效率
Interactivity x Explainability: Toward Understanding How Interactivity Can Improve Computer Vision Explanations
- 通过交互式操作动态调整解释内容
- 用户能更快定位关键信息,理解更深入
- 适合需要深度理解模型决策的科研与应用者
计算机视觉模型的解释对理解模型行为至关重要。然而,现有解释多为静态形式,易导致信息过载、语义与像素层级信息脱节,且缺乏探索空间。本文研究了交互性在三类常见解释方式(基于热力图、概念、原型)中的作用。通过一项包含24名不同技术背景参与者的鸟类识别任务实验,发现交互性虽提升了用户控制力、加速信息定位并深化模型理解,但也带来新挑战。为此,提出设计建议:精心选择默认视图、独立输入控制与受限输出空间。
原文摘要 · Abstract (English)
Explanations for computer vision models are important tools for interpreting how the underlying models work. However, they are often presented in static formats, which pose challenges for users, including information overload, a gap between semantic and pixel-level information, and limited opportunities for exploration. We investigate interactivity as a mechanism for tackling these issues in three common explanation types: heatmap-based, concept-based, and prototype-based explanations. We conducted a study (N=24), using a bird identification task, involving participants with diverse technical and domain expertise. We found that while interactivity enhances user control, facilitates rapid convergence to relevant information, and allows users to expand their understanding of the model and explanation, it also introduces new challenges. To address these, we provide design recommendations for interactive computer vision explanations, including carefully selected default views, independent input controls, and constrained output spaces.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。