arXiv:2504.17921cs.LGcs.AI2025-04ICML被引 10

提出新模型缓解概念模型在分布外时的错误修正失效问题

Avoiding Leakage Poisoning: Concept Interventions Under Distribution Shifts

  • 设计动态机制,仅在数据分布内利用遗漏信息
  • 在分布外输入下,干预后准确率提升显著
  • 适用于需要可解释性与鲁棒性的实际场景

本文研究概念基础模型(CMs)对分布外(OOD)输入的响应。CMs 是可解释的神经架构,先预测一组高层概念(如条纹、黑色),再基于这些概念预测任务标签。我们特别考察了概念干预(即人类专家在测试时纠正模型错误预测的概念)对 OOD 输入下任务预测的影响。分析发现当前先进 CMs 存在一种称为‘泄漏中毒’的缺陷,导致干预后无法有效提升准确率。为此,我们提出 MixCEM,一种能动态利用缺失概念信息的新型 CM,但仅在信息分布内才启用该机制。在有无完整概念标注的任务上,实验表明,MixCEM 在存在或不存在概念干预的情况下,均显著优于强基线,在分布内和分布外样本上都提升了准确率。

原文摘要 · Abstract (English)

In this paper, we investigate how concept-based models (CMs) respond to out-of-distribution (OOD) inputs. CMs are interpretable neural architectures that first predict a set of high-level concepts (e.g., stripes, black) and then predict a task label from those concepts. In particular, we study the impact of concept interventions (i.e., operations where a human expert corrects a CM's mispredicted concepts at test time) on CMs' task predictions when inputs are OOD. Our analysis reveals a weakness in current state-of-the-art CMs, which we term leakage poisoning, that prevents them from properly improving their accuracy when intervened on for OOD inputs. To address this, we introduce MixCEM, a new CM that learns to dynamically exploit leaked information missing from its concepts only when this information is in-distribution. Our results across tasks with and without complete sets of concept annotations demonstrate that MixCEMs outperform strong baselines by significantly improving their accuracy for both in-distribution and OOD samples in the presence and absence of concept interventions.

概念模型分布外可解释性鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。