arXiv:2410.15555cs.LGcs.AI2024-10NeurIPS被引 18

用大模型做概念先验,让可解释模型更准更快

Bayesian Concept Bottleneck Models with LLM Priors

  • 在贝叶斯框架中用LLM自动搜索概念,无需预设候选集
  • 在多类数据上超越可解释模型,甚至逼近黑盒模型性能
  • 能量化不确定性,适合需要可信推理的场景

概念瓶颈模型(CBMs)旨在兼顾可解释性与准确性,但传统方法需预先定义人类可理解的概念集,提取其值并筛选稀疏输入,受限于概念探索范围与标注成本,导致可解释性与准确率权衡。本文提出BC-LLM,通过贝叶斯框架迭代搜索潜在无限的概念集,利用大语言模型(LLM)作为概念提取器和先验知识源。尽管LLM可能误判或虚构信息,本文证明BC-LLM仍可提供严格的统计推断与不确定性量化。在图像、文本和表格数据集上,BC-LLM优于现有可解释基线,部分场景媲美黑盒模型,收敛更快,对分布外样本更具鲁棒性。

原文摘要 · Abstract (English)

Concept Bottleneck Models (CBMs) have been proposed as a compromise between white-box and black-box models, aiming to achieve interpretability without sacrificing accuracy. The standard training procedure for CBMs is to predefine a candidate set of human-interpretable concepts, extract their values from the training data, and identify a sparse subset as inputs to a transparent prediction model. However, such approaches are often hampered by the tradeoff between exploring a sufficiently large set of concepts versus controlling the cost of obtaining concept extractions, resulting in a large interpretability-accuracy tradeoff. This work investigates a novel approach that sidesteps these challenges: BC-LLM iteratively searches over a potentially infinite set of concepts within a Bayesian framework, in which Large Language Models (LLMs) serve as both a concept extraction mechanism and prior. Even though LLMs can be miscalibrated and hallucinate, we prove that BC-LLM can provide rigorous statistical inference and uncertainty quantification. Across image, text, and tabular datasets, BC-LLM outperforms interpretable baselines and even black-box models in certain settings, converges more rapidly towards relevant concepts, and is more robust to out-of-distribution samples.

可解释AI贝叶斯方法大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。