为提示式自然语言解释建立分类框架,助力AI透明治理
A Taxonomy for Design and Evaluation of Prompt-Based Natural Language Explanations
- 从上下文、生成呈现、评估三方面构建提示式解释分类体系
- 提出可系统设计与评估解释的标准化框架
- 适合研究者、审计员及政策制定者参考使用
有效的AI治理需要为利益相关方提供结构化途径以访问和验证AI系统行为。随着大语言模型的兴起,自然语言解释(NLEs)已成为阐明模型行为的关键手段,这要求对其特性与治理影响进行深入分析。本文基于可解释AI(XAI)文献,构建了适用于提示式NLEs的更新版XAI分类体系,涵盖三个维度:(1) 上下文,包括任务、数据、受众和目标;(2) 生成与呈现,涵盖生成方法、输入、交互性、输出形式与表达方式;(3) 评估,关注内容、呈现、用户中心属性及评估环境。该分类体系为研究者、审计员和政策制定者提供了表征、设计与优化NLEs的框架,以支持透明AI系统的建设。
原文摘要 · Abstract (English)
Effective AI governance requires structured approaches for stakeholders to access and verify AI system behavior. With the rise of large language models, Natural Language Explanations (NLEs) are now key to articulating model behavior, which necessitates a focused examination of their characteristics and governance implications. We draw on Explainable AI (XAI) literature to create an updated XAI taxonomy, adapted to prompt-based NLEs, across three dimensions: (1) Context, including task, data, audience, and goals; (2) Generation and Presentation, covering generation methods, inputs, interactivity, outputs, and forms; and (3) Evaluation, focusing on content, presentation, and user-centered properties, as well as the setting of the evaluation. This taxonomy provides a framework for researchers, auditors, and policymakers to characterize, design, and enhance NLEs for transparent AI systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。