arXiv:2412.15628cs.CL2024-12ACL被引 4

测试输入归因能否识别少样本推理中的关键示例

Can Input Attributions Explain Inductive Reasoning in In-Context Learning?

  • 设计心理语言学风格的合成任务,区分关键与模糊示例
  • 简单归因方法表现最佳,大模型更难用梯度法解释
  • 适合研究大模型可解释性与少样本学习机制的学者

理解神经网络内部过程仍是长期挑战,尤其在大语言模型和上下文学习(ICL)时代。例如,如何判断少样本示例中哪个对任务识别或求解起到关键作用,仍不明确。为此,本文设计了一类受心理语言学泛化测试启发的合成诊断任务,其中多数上下文示例在规则上存在歧义,仅一个关键示例能消除歧义。问题在于:传统输入归因(IA)方法能否追踪这一推理过程,即准确识别出该关键示例?实验揭示若干实用发现:特定简单归因方法表现最优,且模型越大,基于梯度的归因方法越难以解释ICL行为。

原文摘要 · Abstract (English)

Interpreting the internal process of neural models has long been a challenge. This challenge remains relevant in the era of large language models (LLMs) and in-context learning (ICL); for example, ICL poses a new issue of interpreting which example in the few-shot examples contributed to identifying/solving the task. To this end, in this paper, we design synthetic diagnostic tasks of inductive reasoning, inspired by the generalization tests typically adopted in psycholinguistics. Here, most in-context examples are ambiguous w.r.t. their underlying rule, and one critical example disambiguates it. The question is whether conventional input attribution (IA) methods can track such a reasoning process, i.e., identify the influential example, in ICL. Our experiments provide several practical findings; for example, a certain simple IA method works the best, and the larger the model, the generally harder it is to interpret the ICL with gradient-based IA methods.

可解释性少样本学习输入归因

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。