arXiv:2503.06730cs.LG2025-03被引 4

用可解释树模型提升概念瓶颈模型的推理能力,支持测试时自适应干预。

Adaptive Test-Time Intervention for Concept Bottleneck Models

  • 用快速可解释贪心求和树生成二值化概念-目标映射关系
  • 在4个数据集上保持原模型性能并实现关键概念识别
  • 适合需要人机协作、有限干预场景的可解释系统

概念瓶颈模型(CBM)通过在深度学习架构中引入人类可理解的‘概念’瓶颈来提升模型可解释性。然而,这些概念如何用于目标预测仍常为黑箱,或为保持可解释性而简化,导致预测性能下降。本文提出使用快速可解释贪心求和树(FIGS)生成二值化蒸馏(BD),将CBM中概念到目标的二值增强部分蒸馏为可解释的树模型,称为FIGS-BD。该方法在保持原CBM教师模型竞争性预测性能的同时,实现了对预测结果的可解释分解,能输出基于二元概念交互的归因,并支持下游任务中的自适应测试时干预。在4个数据集上的实验表明,该方法能有效识别出影响性能的关键概念,在仅允许有限概念干预的人机协同场景中显著提升性能。

原文摘要 · Abstract (English)

Concept bottleneck models (CBM) aim to improve model interpretability by predicting human level "concepts" in a bottleneck within a deep learning model architecture. However, how the predicted concepts are used in predicting the target still either remains black-box or is simplified to maintain interpretability at the cost of prediction performance. We propose to use Fast Interpretable Greedy Sum-Trees (FIGS) to obtain Binary Distillation (BD). This new method, called FIGS-BD, distills a binary-augmented concept-to-target portion of the CBM into an interpretable tree-based model, while maintaining the competitive prediction performance of the CBM teacher. FIGS-BD can be used in downstream tasks to explain and decompose CBM predictions into interpretable binary-concept-interaction attributions and guide adaptive test-time intervention. Across 4 datasets, we demonstrate that our adaptive test-time intervention identifies key concepts that significantly improve performance for realistic human-in-the-loop settings that only allow for limited concept interventions.

可解释性概念瓶颈测试时干预树模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。