arXiv:2609.08585cs.CLcs.LG2026-09

LLM模拟器评估解释效果时会走捷径,可能忽略解释本身。

Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations

论文配图:Limitations of Automated Simulatability: LLM Simulators Can Bypass Explanations
图 1 · 摘自论文原文
  • 用大模型模拟人类预测任务输出,替代人工评估
  • 当类别名有意义时,模拟器直接分类得分高,不依赖解释
  • 匿名类别反而奖励泄露标签映射的解释,适合改进评估设计

可模拟性是一种用于评估解释有用性的方法,通过衡量解释能否帮助用户预测模型输出来量化其价值。由于人工评估成本高,自动化可模拟性使用大语言模型(LLM)作为模拟器替代人类解释接收者,如ConSim(Poché等,2025)所提出,以支持大规模实验。我们定性复现并扩展了ConSim在多个数据集、解释方法族和模拟器大模型上的排名结果,发现了两个关键局限:第一,当类别名称具有语义意义时,模拟器能通过直接解决分类任务获得高可模拟性,无需依赖解释;第二,类别匿名化设置下,某些解释因泄露隐藏标签映射而被错误奖励,我们通过引入新的“类别即概念”基线方法揭示了这一问题。这些结果支持一种捷径假设:在测试场景中,模拟器的预测主要依赖任务先验信息,而解释仅带来微小影响。据此,我们提出改进自动化可模拟性评估的建议。

原文摘要 · Abstract (English)

Simulatability is an evaluation protocol for explanations that quantifies their usefulness by how well they help a user predict a task model's outputs. Since human evaluation is costly, automated simulatability replaces human explainees with LLM simulators, as proposed in ConSim (Poch\'e et al., 2025) for large-scale experiments. We qualitatively replicate and extend ConSim's ranking of explanation methods across the tested datasets, explanation families, and simulator LLMs, and identify two limitations. First, when class names are meaningful, simulators can obtain high simulatability by solving the classification task directly, without relying on the explanations. Second, class anonymization can reward explanations for leaking the hidden label mapping, a limitation we expose with a new classes-as-concepts baseline. These results are consistent with a shortcut hypothesis: in the tested settings, simulator predictions mainly rely on task priors, while explanations produce small changes. We derive recommendations for more robust automated simulatability evaluations.

可模拟性大模型评估解释可信度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。