arXiv:2503.16199cs.LG2025-03NeurIPS被引 9

让可解释模型学会主动求助,提升准确率同时保持透明度。

Deferring Concept Bottleneck Models: Learning to Defer Interventions to Inaccurate Experts

  • 将可解释模型与拒答机制结合,自动判断何时需要人类介入。
  • 在不牺牲解释性前提下,预测性能显著优于传统可解释模型。
  • 适合对决策透明度和可靠性要求高的医疗、金融等场景。

概念瓶颈模型(CBMs)通过基于人类可理解的概念进行预测,提升了模型的可解释性,支持针对性干预。然而,现有方法假设干预者始终存在且能提供正确修正,这在现实中不切实际,因人力成本高且易出错。本文受学习拒答(L2D)启发,提出拒答概念瓶颈模型(DCBMs),使CBMs能够自主学习何时需要人类干预。通过将DCBMs建模为拒答系统的组合,并推导出一致的L2D损失函数进行训练,该框架在保持任务解释性的同时,能说明为何选择拒答。实验表明,尽管需更多依赖人工,但DCBMs在预测性能与可解释性方面均表现优异。

原文摘要 · Abstract (English)

Concept Bottleneck Models (CBMs) are machine learning models that improve interpretability by grounding their predictions on human-understandable concepts, allowing for targeted interventions in their decision-making process. However, when intervened on, CBMs assume the availability of humans that can identify the need to intervene and always provide correct interventions. Both assumptions are unrealistic and impractical, considering labor costs and human error-proneness. In contrast, Learning to Defer (L2D) extends supervised learning by allowing machine learning models to identify cases where a human is more likely to be correct than the model, thus leading to deferring systems with improved performance. In this work, we gain inspiration from L2D and propose Deferring CBMs (DCBMs), a novel framework that allows CBMs to learn when an intervention is needed. To this end, we model DCBMs as a composition of deferring systems and derive a consistent L2D loss to train them. Moreover, by relying on a CBM architecture, DCBMs can explain why defer occurs on the final task. Our results show that DCBMs achieve high predictive performance and interpretability at the cost of deferring more to humans.

可解释模型拒答机制概念瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。