构建首个大模型可解释性标准化评测基准
BELL: Benchmarking the Explainability of Large Language Models
- 提出BELL评测框架,统一评估大模型决策透明度
- 覆盖多种任务,验证主流模型解释能力差异
- 为研究者提供可复现的可解释性分析工具
大语言模型在自然语言处理中展现出卓越能力,但其决策过程往往缺乏透明性,引发信任、偏见和性能评估方面的担忧。为应对这些问题,理解与评估大模型的可解释性至关重要。本文提出了一种标准化的评测方法——大语言模型可解释性基准(Benchmarking the Explainability of Large Language Models, BELL),旨在系统评估大模型的可解释性。该基准涵盖多种自然语言任务,支持对不同模型的解释能力进行量化比较,为提升大模型的可信度和可控性提供关键工具。
原文摘要 · Abstract (English)
Large Language Models have demonstrated remarkable capabilities in natural language processing, yet their decision-making processes often lack transparency. This opaqueness raises significant concerns regarding trust, bias, and model performance. To address these issues, understanding and evaluating the interpretability of LLMs is crucial. This paper introduces a standardised benchmarking technique, Benchmarking the Explainability of Large Language Models, designed to evaluate the explainability of large language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。