arXiv:2505.21740cs.CLcs.AI2025-05被引 5

评估大模型生成任务解释的反事实可模拟性,发现摘要任务效果好,医疗建议仍有不足。

Counterfactual Simulatability of LLM Explanations for Generation Tasks

  • 提出生成任务反事实可模拟性评估框架,用于检验解释是否能预测模型对变体输入的输出
  • 摘要任务中解释帮助用户更准预测反事实输出,医疗建议任务则提升空间大
  • 该评估更适合技能型任务,而非依赖知识的任务,适合关注模型可解释性的研究者

大模型行为具有不可预测性,微小提示变化可能引发意外输出。因此,模型准确解释自身行为的能力至关重要,尤其在高风险场景。一种评估解释有效性的方法是反事实可模拟性——即解释能否帮助用户推断模型在相关反事实输入下的输出。此前该方法仅应用于是非问答任务。本文提出通用框架,将其扩展至生成任务,以新闻摘要和医疗建议为例。结果表明,在摘要任务中,模型解释确实提升了用户对反事实输出的预测能力;但在医疗建议任务中仍存在显著改进空间。此外,研究暗示反事实可模拟性评估更适用于技能型任务,而非知识型任务。

原文摘要 · Abstract (English)

LLMs can be unpredictable, as even slight alterations to the prompt can cause the output to change in unexpected ways. Thus, the ability of models to accurately explain their behavior is critical, especially in high-stakes settings. One approach for evaluating explanations is counterfactual simulatability, how well an explanation allows users to infer the model's output on related counterfactuals. Counterfactual simulatability has been previously studied for yes/no question answering tasks. We provide a general framework for extending this method to generation tasks, using news summarization and medical suggestion as example use cases. We find that while LLM explanations do enable users to better predict LLM outputs on counterfactuals in the summarization setting, there is significant room for improvement for medical suggestion. Furthermore, our results suggest that the evaluation for counterfactual simulatability may be more appropriate for skill-based tasks as opposed to knowledge-based tasks.

可解释性生成任务反事实推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。