用大模型模拟人类评估反事实解释,提升可比性与效率
Towards Unifying Evaluation of Counterfactual Explanations: Leveraging Large Language Models for Human-Centric Assessments
- 用30个场景和206人评分构建评估基准
- 微调LLM后预测准确率达85%(三分类)
- 适合需要真实用户反馈的可解释AI研究者
随着机器学习模型的发展,保持透明性需要更多以人为本的可解释AI方法。反事实解释源于人类推理,能识别使输出发生改变的最小输入变化,对决策支持至关重要。然而,现有评估方法缺乏用户研究基础,指标分散且难以反映人类视角。为此,我们设计了30个反事实场景,收集了206名受访者在8项评估指标上的评分。随后,微调不同大语言模型(LLMs)以预测平均或个体人类判断。该方法使LLMs在零样本评估中最高达到63%准确率,微调后在所有指标上达到85%(三分类)准确率。微调后的模型能更可靠地比较和扩展评估不同反事实解释框架。
原文摘要 · Abstract (English)
As machine learning models evolve, maintaining transparency demands more human-centric explainable AI techniques. Counterfactual explanations, with roots in human reasoning, identify the minimal input changes needed to obtain a given output and, hence, are crucial for supporting decision-making. Despite their importance, the evaluation of these explanations often lacks grounding in user studies and remains fragmented, with existing metrics not fully capturing human perspectives. To address this challenge, we developed a diverse set of 30 counterfactual scenarios and collected ratings across 8 evaluation metrics from 206 respondents. Subsequently, we fine-tuned different Large Language Models (LLMs) to predict average or individual human judgment across these metrics. Our methodology allowed LLMs to achieve an accuracy of up to 63% in zero-shot evaluations and 85% (over a 3-classes prediction) with fine-tuning across all metrics. The fine-tuned models predicting human ratings offer better comparability and scalability in evaluating different counterfactual explanation frameworks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。