用不确定性感知机制提升LLM生成概念标签的可靠性
Uncertainty-aware Language Guidance for Concept Bottleneck Models
- 用分布无关方法量化LLM标注概念的不确定性
- 将不确定性融入训练,提升模型对不可靠标注的鲁棒性
- 适合需高可解释性且依赖LLM构建概念的场景
概念瓶颈模型(CBMs)通过先将输入映射到高层语义概念,再组合这些概念进行分类,具备天然可解释性。然而,人工标注人类可理解的概念需要大量专业知识和人力,限制了CBMs的广泛应用。现有少数工作利用大语言模型(LLMs)构建概念瓶颈,但存在两大局限:一是忽视了LLM标注概念的不确定性,缺乏有效机制来量化这种不确定性,增加了幻觉带来的错误风险;二是未将标注不确定性纳入模型学习过程。为此,我们提出一种新型不确定性感知的CBM方法,不仅能严格量化LLM标注概念标签的不确定性(具备分布无关的保证),还能将量化后的不确定性融入CBM训练中,以应对不同概念标注可靠性的差异。我们还提供了该方法的理论分析。在真实数据集上的大量实验验证了所提方法的有效性。
原文摘要 · Abstract (English)
Concept Bottleneck Models (CBMs) provide inherent interpretability by first mapping input samples to high-level semantic concepts, followed by a combination of these concepts for the final classification. However, the annotation of human-understandable concepts requires extensive expert knowledge and labor, constraining the broad adoption of CBMs. On the other hand, there are a few works that leverage the knowledge of large language models (LLMs) to construct concept bottlenecks. Nevertheless, they face two essential limitations: First, they overlook the uncertainty associated with the concepts annotated by LLMs and lack a valid mechanism to quantify uncertainty about the annotated concepts, increasing the risk of errors due to hallucinations from LLMs. Additionally, they fail to incorporate the uncertainty associated with these annotations into the learning process for concept bottleneck models. To address these limitations, we propose a novel uncertainty-aware CBM method, which not only rigorously quantifies the uncertainty of LLM-annotated concept labels with valid and distribution-free guarantees, but also incorporates quantified concept uncertainty into the CBM training procedure to account for varying levels of reliability across LLM-annotated concepts. We also provide the theoretical analysis for our proposed method. Extensive experiments on the real-world datasets validate the desired properties of our proposed methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。