arXiv:2412.14097cs.LGcs.AI2024-12被引 12

让大模型在数据变化时仍能解释决策,且准确率提升28%。

Adaptive Concept Bottleneck for Foundation Models Under Distribution Shifts

  • 用可解释的概念向量动态调整大模型推理过程。
  • 仅用目标域无标签数据,使解释更贴合实际测试数据。
  • 适合医疗、金融等需可信决策的高风险场景。

基础模型(FMs)的进展推动了机器学习范式变革。这些预训练的大模型通过浅层全连接网络进行轻量微调,广泛应用于下游任务。然而其黑箱特性在医疗、金融、安全等关键领域带来挑战。本文探索使用概念瓶颈模型(CBM)将复杂非可解释的模型转化为基于高层概念向量的可解释决策流程。特别关注测试阶段在真实环境中的部署,此时输入分布常发生偏移。我们识别了不同分布偏移下的潜在失效模式,并提出自适应概念瓶颈框架,仅依赖目标域的无标签数据,动态调整概念向量库和预测层,无需源数据。在多种真实分布偏移上的实验表明,该方法生成的概念解释更符合测试数据,后部署准确率最高提升28%,使CBM性能逼近不可解释分类器水平。

原文摘要 · Abstract (English)

Advancements in foundation models (FMs) have led to a paradigm shift in machine learning. The rich, expressive feature representations from these pre-trained, large-scale FMs are leveraged for multiple downstream tasks, usually via lightweight fine-tuning of a shallow fully-connected network following the representation. However, the non-interpretable, black-box nature of this prediction pipeline can be a challenge, especially in critical domains such as healthcare, finance, and security. In this paper, we explore the potential of Concept Bottleneck Models (CBMs) for transforming complex, non-interpretable foundation models into interpretable decision-making pipelines using high-level concept vectors. Specifically, we focus on the test-time deployment of such an interpretable CBM pipeline "in the wild", where the input distribution often shifts from the original training distribution. We first identify the potential failure modes of such a pipeline under different types of distribution shifts. Then we propose an adaptive concept bottleneck framework to address these failure modes, that dynamically adapts the concept-vector bank and the prediction layer based solely on unlabeled data from the target domain, without access to the source (training) dataset. Empirical evaluations with various real-world distribution shifts show that our adaptation method produces concept-based interpretations better aligned with the test data and boosts post-deployment accuracy by up to 28%, aligning the CBM performance with that of non-interpretable classification.

可解释性大模型分布偏移概念瓶颈

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。