让大模型用普通人能懂的方式解释自己,提升高风险领域可信度。
Can LLMs faithfully generate their layperson-understandable 'self'?: A Case Study in High-Stakes Domains
- 提出'ReQuesting'新方法,让大模型生成可理解的解释。
- 在法律、医疗、金融三领域实现高可复现的通俗解释。
- 解释内容与模型内在推理高度一致,适合非专业人士使用。
大型语言模型(LLMs)已深刻影响人类知识的各个领域。然而,这些模型对普通人的可解释性——对建立信任至关重要——仍受到多种质疑。本文在法律、医疗和金融三个高优先级应用领域,引入一种面向普通人的新型可解释性概念,称为‘ReQuesting’,并使用多个最先进的大模型进行验证。该方法在多个任务上实现了高可复现性的、可理解的通俗化解释生成。此外,我们观察到这些可解释算法与大模型自身的内在推理过程存在显著一致性。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have significantly impacted nearly every domain of human knowledge. However, the explainability of these models esp. to laypersons, which are crucial for instilling trust, have been examined through various skeptical lenses. In this paper, we introduce a novel notion of LLM explainability to laypersons, termed $\textit{ReQuesting}$, across three high-priority application domains -- law, health and finance, using multiple state-of-the-art LLMs. The proposed notion exhibits faithful generation of explainable layman-understandable algorithms on multiple tasks through high degree of reproducibility. Furthermore, we observe a notable alignment of the explainable algorithms with intrinsic reasoning of the LLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。