用AI生成的虚拟人物评估解释性,让XAI选型更懂用户
VirtualXAI: A User-Centric Framework for Explainability Assessment Leveraging GPT-Generated Personas
- 用大模型生成的虚拟人物模拟真实用户,评估XAI方法的可理解性
- 结合数据特征推荐最佳模型与解释方法,提升人机协作效率
- 既看数字指标又看用户体验,解决传统评估重定量轻定性的缺陷
在数据驱动时代,人工智能在产业数字化中扮演关键角色。当前对可解释AI(XAI)的需求上升,旨在提升模型的可解释性、透明度和可信度。然而,现有评估框架多聚焦于保真度、一致性、稳定性等定量指标,忽视了满意度、可理解性等定性特征。同时,从业者在选择数据集、AI模型和XAI方法时缺乏有效指导,制约人机协同。为此,我们提出一个以用户为中心的评估框架VirtualXAI,利用大语言模型生成的“人物背景故事库”构建虚拟用户,融合定量基准测试与定性用户评估。该框架还引入基于内容的推荐系统,根据输入数据特征从基准数据集库中匹配最优方案,输出预估的XAI评分,并为特定场景推荐最适配的AI模型与XAI方法。
原文摘要 · Abstract (English)
In today's data-driven era, computational systems generate vast amounts of data that drive the digital transformation of industries, where Artificial Intelligence (AI) plays a key role. Currently, the demand for eXplainable AI (XAI) has increased to enhance the interpretability, transparency, and trustworthiness of AI models. However, evaluating XAI methods remains challenging: existing evaluation frameworks typically focus on quantitative properties such as fidelity, consistency, and stability without taking into account qualitative characteristics such as satisfaction and interpretability. In addition, practitioners face a lack of guidance in selecting appropriate datasets, AI models, and XAI methods -a major hurdle in human-AI collaboration. To address these gaps, we propose a framework that integrates quantitative benchmarking with qualitative user assessments through virtual personas based on the "Anthology" of backstories of the Large Language Model (LLM). Our framework also incorporates a content-based recommender system that leverages dataset-specific characteristics to match new input data with a repository of benchmarked datasets. This yields an estimated XAI score and provides tailored recommendations for both the optimal AI model and the XAI method for a given scenario.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。