提升概念模型可靠性,让机器决策更可信。
Towards Robust and Reliable Concept Representations: Reliability-Enhanced Concept Embedding Model
- 分离无关特征,确保概念表征聚焦关键属性
- 跨样本语义对齐,解决同一概念表达不一致问题
- 在背景和领域变化下表现更稳定,适合高可靠性场景
概念瓶颈模型(CBMs)通过预测人类可理解的概念作为决策中间步骤来增强可解释性。然而,这些模型常因概念表征不可靠而影响下游任务的鲁棒性,尤其在分布偏移时更为明显。其核心问题在于对无关特征(如背景变化)敏感,以及同一概念在不同样本间缺乏语义一致性。为此,我们提出可靠性增强的概念嵌入模型(RECEM),采用双策略:概念级解耦以分离无关特征,概念混合机制确保跨样本语义对齐。二者协同提升概念可靠性,使模型专注有意义的对象属性并生成忠实的概念表示。实验表明,RECEM在多个数据集上持续优于现有基线,在背景和领域偏移下表现更优,验证了解耦与对齐策略在提升CBMs可靠性与鲁棒性方面的有效性。
原文摘要 · Abstract (English)
Concept Bottleneck Models (CBMs) aim to enhance interpretability by predicting human-understandable concepts as intermediates for decision-making. However, these models often face challenges in ensuring reliable concept representations, which can propagate to downstream tasks and undermine robustness, especially under distribution shifts. Two inherent issues contribute to concept unreliability: sensitivity to concept-irrelevant features (e.g., background variations) and lack of semantic consistency for the same concept across different samples. To address these limitations, we propose the Reliability-Enhanced Concept Embedding Model (RECEM), which introduces a two-fold strategy: Concept-Level Disentanglement to separate irrelevant features from concept-relevant information and a Concept Mixup mechanism to ensure semantic alignment across samples. These mechanisms work together to improve concept reliability, enabling the model to focus on meaningful object attributes and generate faithful concept representations. Experimental results demonstrate that RECEM consistently outperforms existing baselines across multiple datasets, showing superior performance under background and domain shifts. These findings highlight the effectiveness of disentanglement and alignment strategies in enhancing both reliability and robustness in CBMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。