自解释能提升对大模型行为的预测能力,且优于外部模型生成的解释。
A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
- 提出新度量NSG,通过预测能力评估自解释是否忠实
- 自解释使行为预测准确率提升11%-37%,优于外部模型解释
- 发现5%-15%自解释严重误导,但整体仍具预测价值
大模型自解释常被视为人工智能监管的潜力工具,但其与真实推理过程的一致性尚不明确。现有可信度评估方法多依赖对抗提示或错误检测,忽视解释的预测价值。本文提出归一化可模拟增益(NSG),基于忠实解释应使观察者掌握模型决策标准、从而更好预测其在相关输入上的行为这一理念。我们在7,000个来自健康、商业和伦理领域的反事实样本上,评估了18个前沿专有及开源模型(如Gemini 3、GPT-5.2、Claude 4.5)。结果表明,自解释显著提升模型行为预测能力,NSG提升11%-37%;即使外部更强模型生成的解释,也难以超越自解释的预测信息量。这表明自知识具有不可替代优势。同时发现,跨模型中5%-15%的自解释存在严重误导。尽管不完美,本研究为自解释提供了积极证据:它们确实编码了有助于预测模型行为的信息。
原文摘要 · Abstract (English)
LLM self-explanations are often presented as a promising tool for AI oversight, yet their faithfulness to the model's true reasoning process is poorly understood. Existing faithfulness metrics have critical limitations, typically relying on identifying unfaithfulness via adversarial prompting or detecting reasoning errors. These methods overlook the predictive value of explanations. We introduce Normalized Simulatability Gain (NSG), a general and scalable metric based on the idea that a faithful explanation should allow an observer to learn a model's decision-making criteria, and thus better predict its behavior on related inputs. We evaluate 18 frontier proprietary and open-weight models, e.g., Gemini 3, GPT-5.2, and Claude 4.5, on 7,000 counterfactuals from popular datasets covering health, business, and ethics. We find self-explanations substantially improve prediction of model behavior (11-37% NSG). Self-explanations also provide more predictive information than explanations generated by external models, even when those models are stronger. This implies an advantage from self-knowledge that external explanation methods cannot replicate. Our approach also reveals that, across models, 5-15% of self-explanations are egregiously misleading. Despite their imperfections, we show a positive case for self-explanations: they encode information that helps predict model behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。