arXiv:2501.11827cs.LGcs.AI2025-01

为生成模型设计可定制的后验解释方法,提升透明度与可信度。

PXGen: A Post-hoc Explainable Method for Generative Models

  • 通过锚集与内外部标准生成可定制解释材料
  • 基于特征值提供示例化解释并可视化展示
  • 适合关注生成模型可信性与可解释性的研究者

随着生成式AI在众多应用中的快速发展,可解释AI(XAI)在确保其负责任开发与部署中发挥关键作用。近年来,XAI在透明性、可解释性和可信度方面取得显著进展。高效XAI方法需满足两大核心标准:一是解释的质量与流畅性,包括忠实性、合理性、完整性及个性化适配;二是系统设计的可靠性、鲁棒性、输出可验证性与算法透明性。然而,针对生成模型的XAI研究仍相对匮乏,缺乏有效方法满足上述标准。本文提出PXGen,一种面向生成模型的后验可解释方法。给定待解释模型,PXGen准备两部分材料:锚集与内在/外在评估标准,用户可根据需求自定义。通过计算各标准下的特征值,对每个锚点生成一组特征值,并依据这些值采用基于示例的解释方法,结合如k-分散或k-中心等可追踪算法进行可视化呈现。

原文摘要 · Abstract (English)

With the rapid growth of generative AI in numerous applications, explainable AI (XAI) plays a crucial role in ensuring the responsible development and deployment of generative AI technologies. XAI has undergone notable advancements and widespread adoption in recent years, reflecting a concerted push to enhance the transparency, interpretability, and credibility of AI systems. Recent research emphasizes that a proficient XAI method should adhere to a set of criteria, primarily focusing on two key areas. Firstly, it should ensure the quality and fluidity of explanations, encompassing aspects like faithfulness, plausibility, completeness, and tailoring to individual needs. Secondly, the design principle of the XAI system or mechanism should cover the following factors such as reliability, resilience, the verifiability of its outputs, and the transparency of its algorithm. However, research in XAI for generative models remains relatively scarce, with little exploration into how such methods can effectively meet these criteria in that domain. In this work, we propose PXGen, a post-hoc explainable method for generative models. Given a model that needs to be explained, PXGen prepares two materials for the explanation, the Anchor set and intrinsic & extrinsic criteria. Those materials are customizable by users according to their purpose and requirements. Via the calculation of each criterion, each anchor has a set of feature values and PXGen provides examplebased explanation methods according to the feature values among all the anchors and illustrated and visualized to the users via tractable algorithms such as k-dispersion or k-center.

可解释AI生成模型后验解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。