arXiv:2502.20635cs.HCcs.LG2025-02被引 3

用大模型评估机器学习解释质量,发现其能辅助但无法替代人类判断。

Can LLM Assist in the Evaluation of the Quality of Machine Learning Explanations?

  • 设计混合评估流程,结合大模型与人工评委
  • 大模型在主观评价上表现良好,但不如人类可靠
  • 适合需要快速初筛解释质量的研究场景

可解释机器学习(XML)旨在揭示机器学习系统‘黑箱’结果的机制。尽管已有多种解释方法,但在特定场景下选择最优方法仍不明确,亟需有效评估手段。基于Transformer的大语言模型(LLM)具备评估能力,为采用LLM作为评判者提供了可能。本文提出一种融合LLM与人工评委的评估工作流,在鸢尾花分类任务中,通过主观与客观指标比较了LLM评委与人类评委的表现。结果表明,虽然LLM在主观评价上能有效评估解释质量,但尚未达到可完全替代人类评委的水平。

原文摘要 · Abstract (English)

EXplainable machine learning (XML) has recently emerged to address the mystery mechanisms of machine learning (ML) systems by interpreting their 'black box' results. Despite the development of various explanation methods, determining the most suitable XML method for specific ML contexts remains unclear, highlighting the need for effective evaluation of explanations. The evaluating capabilities of the Transformer-based large language model (LLM) present an opportunity to adopt LLM-as-a-Judge for assessing explanations. In this paper, we propose a workflow that integrates both LLM-based and human judges for evaluating explanations. We examine how LLM-based judges evaluate the quality of various explanation methods and compare their evaluation capabilities to those of human judges within an iris classification scenario, employing both subjective and objective metrics. We conclude that while LLM-based judges effectively assess the quality of explanations using subjective metrics, they are not yet sufficiently developed to replace human judges in this role.

可解释性大模型评估人机协同

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。