arXiv:2602.02304cs.AIcs.LG2026-02

提出新解释框架,专门分析大模型行为变化的因果机制。

Comparing Explanations is Not Enough, Explain the Change: New Standards are Needed to Explain Behavioral Shifts in Large Language Models

  • 构建对比XAI(XAIΔ),聚焦干预前后模型行为的差异
  • 验证该方法可生成可比、可操作且可监控的解释结果
  • 适用于监管审计与重大系统变更的合规记录

大规模基础模型在缩放、微调、基于人类反馈的强化学习或上下文学习等干预下会表现出行为变化。当前可解释性方法因将模型视为静态对象,或仅比较不同检查点的独立解释,无法有效解释行为转变过程,造成欧盟《人工智能法案》、美国各州立法及中国AI监管中对重大系统修改需追溯因果链的要求落空。本文主张以行为变化本身为解释核心,提出新型对比XAI(XAI$_Δ$)范式,并定义可比性、有效性、可行动性与可监测性四项标准,旨在使模型审计建立在明确可测的基础上。通过初步实验验证其必要性,生成可用于治理与事件记录的过渡报告。

原文摘要 · Abstract (English)

Large-scale foundation models exhibit behavioral shifts when subjected to interventions such as scaling, fine-tuning, reinforcement learning with human feedback, or in-context learning. Current explainability methods are structurally ill-suited to explain these shifts, because they either treat models as static objects, as traditional eXplainable AI (XAI) approaches do, or merely compare independent explanations across different checkpoints of a model. As a result, these approaches fail to explain the functional transition between two model instances in which a certain behavior has shifted following an intervention. This gap creates significant governance risks across jurisdictions including the EU AI Act, US state legislation, and Chinese AI regulations, which require documenting causal chains for substantial system modifications. This position paper argues that explaining behavioral shifts in large language models requires a principled approach that treats the shift itself as the primary object of explanation: namely, one that explains how and why an intervention transforms a reference model into an updated model with different behavior. To support this claim, we introduce Comparative XAI (XAI$_Δ$), a novel XAI paradigm aimed at explaining the difference between two model checkpoints where a behavior has shifted, together with a set of desiderata specifying what XAI$_Δ$ explainers and explanations must satisfy, including comparability, validity, actionability, and monitoring, with the goal of grounding model auditing in explicit, measurable requirements. Finally, we provide preliminary evidence suggesting the need for XAI$_Δ$ in practice through illustrative experiments, compiling the resulting findings into a transition report directly usable for governance and incident documentation.

可解释性大模型监管行为变化XAI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。