用自动化指标评估大模型生成的可解释性叙事质量
How good is my story? Towards quantitative metrics for evaluating LLM-generated XAI narratives
- 提出框架与多个自动评估指标,替代人工评测
- 在多个数据集和提示类型上对比主流大模型表现
- 发现大模型在可解释性叙述中存在幻觉新问题
大模型在可解释人工智能(XAI)中的一个快速发展应用是将如SHAP等定量解释转化为用户友好的叙述,以解释小型预测模型的决策。在此领域,无需依赖人类偏好研究或调查来评估这些叙述正变得日益重要。本文提出一个框架,并探索多种自动化指标,用于评估大模型在表格分类任务中生成的解释性叙述。我们利用该方法在不同数据集和提示类型下比较几种最先进的大模型。作为其应用示范,这些指标帮助我们识别出大模型在生成可解释性叙述时面临的新挑战,尤其是幻觉问题。
原文摘要 · Abstract (English)
A rapidly developing application of LLMs in XAI is to convert quantitative explanations such as SHAP into user-friendly narratives to explain the decisions made by smaller prediction models. Evaluating the narratives without relying on human preference studies or surveys is becoming increasingly important in this field. In this work we propose a framework and explore several automated metrics to evaluate LLM-generated narratives for explanations of tabular classification tasks. We apply our approach to compare several state-of-the-art LLMs across different datasets and prompt types. As a demonstration of their utility, these metrics allow us to identify new challenges related to LLM hallucinations for XAI narratives.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。