arXiv:2410.17781cs.AI2024-10被引 13

用大模型模拟用户,低成本高效评估AI解释效果

Evaluating Explanations Through LLMs: Beyond Traditional User Studies

  • 用7个大模型模拟人类参与者,复现解释效果评估实验
  • 大模型能复现原始研究的大部分结论,但表现因模型而异
  • 模型记忆和输出波动会影响结果与真人的一致性

随着人工智能在医疗等关键领域广泛应用,可解释AI(XAI)工具对建立信任与透明度至关重要。然而,传统的人类用户研究评估方式成本高、耗时长且难扩展。本文探索使用大语言模型(LLMs)模拟真实参与者,以简化XAI评估流程。我们复现了一项对比反事实与因果解释的用户研究,利用七种LLM在不同设置下模拟人类参与者的反应。结果显示:(i) LLMs能有效复现原研究的多数结论;(ii) 不同LLM的结果一致性存在差异;(iii) 实验因素如模型记忆能力与输出变异性显著影响其与人类响应的一致性。初步结果表明,LLMs可为定性XAI评估提供一种可扩展且经济高效的替代方案。

原文摘要 · Abstract (English)

As AI becomes fundamental in sectors like healthcare, explainable AI (XAI) tools are essential for trust and transparency. However, traditional user studies used to evaluate these tools are often costly, time consuming, and difficult to scale. In this paper, we explore the use of Large Language Models (LLMs) to replicate human participants to help streamline XAI evaluation. We reproduce a user study comparing counterfactual and causal explanations, replicating human participants with seven LLMs under various settings. Our results show that (i) LLMs can replicate most conclusions from the original study, (ii) different LLMs yield varying levels of alignment in the results, and (iii) experimental factors such as LLM memory and output variability affect alignment with human responses. These initial findings suggest that LLMs could provide a scalable and cost-effective way to simplify qualitative XAI evaluation.

可解释AI大模型评估用户研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。