arXiv:2601.03156cs.LGcs.AI2026-01

通过反事实提示解释,揭示大模型输出偏见的根源。

Prompt-Counterfactual Explanations for Generative AI System Behavior

  • 用反事实提示法分析输入如何导致模型输出特定特征
  • 可识别引发毒性、政治倾向等的敏感提示词
  • 适合用于模型安全测试与合规性审查

随着生成式AI系统在现实场景中的广泛应用,组织亟需理解其行为机制。本文聚焦关键问题:大语言模型生成特定输出特征(如毒性、负面情绪或政治偏见)的根源是否来自输入提示?传统反事实解释不适用于非确定性的生成式AI系统,因此我们提出一种灵活框架,结合下游分类器识别输出特征,并设计算法生成提示反事实解释(PCE)。通过三个案例研究,验证了PCE在抑制不良输出和增强红队测试方面的有效性。该方法为生成式AI的提示层面可解释性奠定基础,对高风险应用及透明度监管具有重要意义。

原文摘要 · Abstract (English)

As generative AI systems become integrated into real-world applications, organizations increasingly need to be able to understand and interpret their behavior. In particular, decision-makers need to understand what causes generative AI systems to exhibit specific output characteristics. Within this general topic, this paper examines a key question: what is it about the input -- the prompt -- that causes an LLM-based generative AI system to produce output that exhibits specific characteristics, such as toxicity, negative sentiment, or political bias. To examine this question, we adapt a common technique from the Explainable AI literature: counterfactual explanations. We explain why traditional counterfactual explanations cannot be applied directly to generative AI systems, due to several differences in how generative AI systems function. We then propose a flexible framework that adapts counterfactual explanations to non-deterministic, generative AI systems in scenarios where downstream classifiers can reveal key characteristics of their outputs. Based on this framework, we introduce an algorithm for generating prompt-counterfactual explanations (PCEs). Finally, we demonstrate the production of counterfactual explanations for generative AI systems with three case studies, examining different output characteristics (viz., political leaning, toxicity, and sentiment). The case studies further show that PCEs can streamline prompt engineering to suppress undesirable output characteristics and can enhance red-teaming efforts to uncover additional prompts that elicit undesirable outputs. Ultimately, this work lays a foundation for prompt-focused interpretability in generative AI: a capability that will become indispensable as these models are entrusted with higher-stakes tasks and subject to emerging regulatory requirements for transparency and accountability.

可解释性提示工程AI安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。