arXiv:2601.12049cs.CVcs.AI2026-01被引 1

用逻辑表达式解析视觉模型决策依据,提升可解释性。

\textit{FocaLogic}: Logic-Based Interpretation of Visual Model Decisions

  • 通过识别关键视觉区域生成逻辑表达式解释模型判断。
  • 提出精度、召回率等量化指标评估模型行为。
  • 适合需要透明决策的医疗、自动驾驶等高风险场景。

现代视觉模型的可解释性在高风险应用中至关重要。现有方法或依赖白盒模型访问,或缺乏定量严谨性。为此,我们提出FocaLogic,一种无需模型内部信息的新型框架,通过逻辑表示来解释和量化视觉模型的决策过程。FocaLogic识别影响预测的关键视觉区域(称为视觉焦点),并将其转化为精确简洁的逻辑表达式,实现透明且结构化的解释。我们还设计了一套量化指标,包括焦点精度、召回率和差异度,以客观评估模型在多种场景下的表现。实证分析表明,FocaLogic能揭示训练引发的关注集中、泛化带来的焦点准确率提升,以及偏见和对抗攻击下的异常关注现象。总体而言,FocaLogic提供了一种系统、可扩展且量化的视觉模型解释方案。

原文摘要 · Abstract (English)

Interpretability of modern visual models is crucial, particularly in high-stakes applications. However, existing interpretability methods typically suffer from either reliance on white-box model access or insufficient quantitative rigor. To address these limitations, we introduce FocaLogic, a novel model-agnostic framework designed to interpret and quantify visual model decision-making through logic-based representations. FocaLogic identifies minimal interpretable subsets of visual regions-termed visual focuses-that decisively influence model predictions. It translates these visual focuses into precise and compact logical expressions, enabling transparent and structured interpretations. Additionally, we propose a suite of quantitative metrics, including focus precision, recall, and divergence, to objectively evaluate model behavior across diverse scenarios. Empirical analyses demonstrate FocaLogic's capability to uncover critical insights such as training-induced concentration, increasing focus accuracy through generalization, and anomalous focuses under biases and adversarial attacks. Overall, FocaLogic provides a systematic, scalable, and quantitative solution for interpreting visual models.

可解释性逻辑表达视觉模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。