arXiv:2512.17891cs.CV2025-12

无需训练即可让ViT模型自解释,靠关键点匹配实现可视化决策

Keypoint Counting Classifiers: Turning Vision Transformers into Self-Explainable Models Without Training

  • 利用ViT自动识别图像关键点,构建可解释的分类流程
  • 相比基线方法显著提升人机沟通效果,无需重新训练
  • 适合希望快速提升ViT模型透明度的研究者与工程师

当前自解释模型(SEMs)的设计依赖复杂的训练过程和特定架构,难以实际应用。随着基于视觉变压器(ViTs)的通用基础模型发展,这一问题愈发突出。因此亟需新方法为ViT类基础模型提供透明性与可靠性。本文提出一种无需重训练即可将任意已训练好的ViT模型转化为自解释模型的方法,称为关键点计数分类器(KCCs)。近期研究表明,ViTs能高精度自动识别图像间的匹配关键点,我们在此基础上构建了可直观可视化的决策机制。大量实验表明,相较于现有基线方法,KCCs显著提升了人机交互效果。我们认为,KCCs是使ViT类基础模型更透明、更可靠的重要一步。

原文摘要 · Abstract (English)

Current approaches for designing self-explainable models (SEMs) require complicated training procedures and specific architectures which makes them impractical. With the advance of general purpose foundation models based on Vision Transformers (ViTs), this impracticability becomes even more problematic. Therefore, new methods are necessary to provide transparency and reliability to ViT-based foundation models. In this work, we present a new method for turning any well-trained ViT-based model into a SEM without retraining, which we call Keypoint Counting Classifiers (KCCs). Recent works have shown that ViTs can automatically identify matching keypoints between images with high precision, and we build on these results to create an easily interpretable decision process that is inherently visualizable in the input. We perform an extensive evaluation which show that KCCs improve the human-machine communication compared to recent baselines. We believe that KCCs constitute an important step towards making ViT-based foundation models more transparent and reliable.

自解释ViT可解释性关键点

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。