arXiv:2602.01206cs.AI2026-02

提出可解释生成式AI的新方法,让模型决策过程透明可控。

Addressing Explainability of Generative AI using SMILE (Statistical Model-agnostic Interpretability with Local Explanations)

  • 用输入扰动+距离度量+代理模型,量化提示词各部分对输出影响。
  • 在大模型上生成细粒度词级归因热图,准确追踪推理路径。
  • 适用于文本与图像生成模型,适合高风险场景的可信部署评估。

生成式人工智能快速发展,能够生成复杂文本与视觉内容,但其决策过程仍高度不透明,限制了在高风险场景中的信任与问责。本文提出gSMILE,一种统一的生成式模型可解释性框架,将统计模型无关的局部解释方法(SMILE)扩展至生成场景。gSMILE通过受控的文本输入扰动、Wasserstein距离度量和加权代理建模,量化并可视化提示词或指令中特定成分对模型输出的影响。应用于大语言模型(LLMs)时,可实现细粒度的词级别归因,并生成直观的热图以突出关键词与推理路径;在基于指令的图像编辑模型中,采用精确的文本扰动机制,分析编辑指令修改如何影响生成图像。结合基于操作设计域(ODD)框架的情境化评估策略,gSMILE可系统评估模型在多样化语义与环境条件下的行为表现。为评估解释质量,定义了稳定性、保真度、准确性、一致性与忠实性等严格指标,并在多种生成架构上进行验证。大量实验表明,gSMILE能生成稳健且符合人类认知的归因结果,并有效泛化至当前主流生成模型。这些发现凸显了gSMILE在推动生成式AI技术透明、可靠与负责任部署方面的潜力。

原文摘要 · Abstract (English)

The rapid advancement of generative artificial intelligence has enabled models capable of producing complex textual and visual outputs; however, their decision-making processes remain largely opaque, limiting trust and accountability in high-stakes applications. This thesis introduces gSMILE, a unified framework for the explainability of generative models, extending the Statistical Model-agnostic Interpretability with Local Explanations (SMILE) method to generative settings. gSMILE employs controlled perturbations of textual input, Wasserstein distance metrics, and weighted surrogate modelling to quantify and visualise how specific components of a prompt or instruction influence model outputs. Applied to Large Language Models (LLMs), gSMILE provides fine-grained token-level attribution and generates intuitive heatmaps that highlight influential tokens and reasoning pathways. In instruction-based image editing models, the exact text-perturbation mechanism is employed, allowing for the analysis of how modifications to an editing instruction impact the resulting image. Combined with a scenario-based evaluation strategy grounded in the Operational Design Domain (ODD) framework, gSMILE allows systematic assessment of model behaviour across diverse semantic and environmental conditions. To evaluate explanation quality, we define rigorous attribution metrics, including stability, fidelity, accuracy, consistency, and faithfulness, and apply them across multiple generative architectures. Extensive experiments demonstrate that gSMILE produces robust, human-aligned attributions and generalises effectively across state-of-the-art generative models. These findings highlight the potential of gSMILE to advance transparent, reliable, and responsible deployment of generative AI technologies.

可解释性生成模型大模型归因分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。