arXiv:2606.01148cs.CL2026-06

比较两种模型解释方式的可模拟性,发现不同解释格式影响理解效果。

Not All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales

论文配图:Not All Explanations Simulate Equally: Comparing Verbalized Feature Attributions and Self-Generated Rationales
图 1 · 摘自论文原文
  • 用反事实测试评估两种解释的可模拟性:特征归因与自生成推理。
  • 自生成推理在多数情况下更利于预测模型对后续问题的回答。
  • 解释格式和粒度显著影响人类或模型能否准确模拟原模型行为。

自然语言解释常被视为理解模型行为的统一接口,但不同来源的解释在支持模拟方面存在差异。本文比较了问答模型的两类解释:语义化特征归因与自生成推理。我们在统一的反事实模拟设置下进行评估,使用大语言模型作为裁判,测量其能否根据解释更准确地预测模型对后续问题的回答。针对多个指令微调模型,分析解释来源、表述策略及特征粒度对解释可模拟性的影响。结果表明,解释格式与粒度显著影响可模拟性:基于归因的解释与自生成推理在提升反事实预测能力方面表现不同,且这种差异在不同模型和格式间存在变化。

原文摘要 · Abstract (English)

Natural-language explanations are often treated as a unified interface for understanding model behavior, but different explanation sources may support simulation in different ways. This paper compares two families of explanations for question answering models: verbalized feature attributions and self-generated rationales. We evaluate them under a shared counterfactual simulation setting, using an LLM judge as predictor and measuring whether it can better predict a model's answers to follow-up questions when given its explanation. Across multiple instruction-tuned models, we analyze how explanation source, verbalization strategy, and feature granularity affect the simulatability of explanations. Our results show that explanation format and granularity affect simulatability: attribution-based explanations and self-generated rationales differ in how much they improve counterfactual prediction, with effects that vary across models and formats.

模型解释可模拟性自然语言推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。