构建可信赖的解释评估框架,用结构化反事实数据测试大模型概念解释的准确性。
LIBERTy: A Causal Framework for Benchmarking Concept-Based Explanations of LLMs with Structural Counterfactuals
- 基于结构因果模型生成可计算的反事实样本对
- 在3个真实场景中验证解释方法,发现现有方法仍有巨大提升空间
- 揭示闭源模型对人口属性干预不敏感,可能因后期训练缓解偏见
概念解释通过量化性别、经验等高层概念对模型行为的影响,对高风险领域决策至关重要。现有评估依赖人工编写的反事实,成本高且不准确。为此,我们提出LIBERTy框架,基于显式定义的文本生成结构因果模型(SCM),通过干预概念并传播至大模型生成反事实输出,构建结构化反事实对。我们构建了疾病检测、简历筛选和职场暴力预测三个数据集,并引入新评估指标“顺序忠实性”。在五种模型上评估多种方法后发现,概念解释仍存在显著改进空间。该框架还可系统分析模型对干预的敏感性:发现闭源大模型对人口概念干预反应明显减弱,可能源于后期训练中的偏见缓解措施。LIBERTy为开发可信解释方法提供了关键基准。
原文摘要 · Abstract (English)
Concept-based explanations quantify how high-level concepts (e.g., gender or experience) influence model behavior, which is crucial for decision-makers in high-stakes domains. Recent work evaluates the faithfulness of such explanations by comparing them to reference causal effects estimated from counterfactuals. In practice, existing benchmarks rely on costly human-written counterfactuals that serve as an imperfect proxy. To address this, we introduce a framework for constructing datasets containing structural counterfactual pairs: LIBERTy (LLM-based Interventional Benchmark for Explainability with Reference Targets). LIBERTy is grounded in explicitly defined Structured Causal Models (SCMs) of the text generation, interventions on a concept propagate through the SCM until an LLM generates the counterfactual. We introduce three datasets (disease detection, CV screening, and workplace violence prediction) together with a new evaluation metric, order-faithfulness. Using them, we evaluate a wide range of methods across five models and identify substantial headroom for improving concept-based explanations. LIBERTy also enables systematic analysis of model sensitivity to interventions: we find that proprietary LLMs show markedly reduced sensitivity to demographic concepts, likely due to post-training mitigation. Overall, LIBERTy provides a much-needed benchmark for developing faithful explainability methods.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。