arXiv:2411.02018cs.CLcs.AI2024-11综述被引 12

剖析大模型在上下文学习中的捷径现象及其应对策略。

Shortcut Learning in In-Context Learning: A Survey

  • 系统梳理上下文学习中出现的各类决策捷径
  • 总结现有基准与缓解捷径的学习策略
  • 适合关注大模型泛化能力的研究者阅读

捷径学习指模型在实际任务中采用简单、非鲁棒的决策规则,影响其泛化与鲁棒性。近年来大语言模型快速发展,越来越多研究揭示了捷径学习对大模型的影响。本文从新视角综述上下文学习(ICL)中捷径学习的相关研究,深入探讨了ICL任务中的捷径类型、成因、可用基准及缓解策略。基于观察,总结现有研究未解问题,并尝试勾勒未来研究方向。

原文摘要 · Abstract (English)

Shortcut learning refers to the phenomenon where models employ simple, non-robust decision rules in practical tasks, which hinders their generalization and robustness. With the rapid development of large language models (LLMs) in recent years, an increasing number of studies have shown the impact of shortcut learning on LLMs. This paper provides a novel perspective to review relevant research on shortcut learning in In-Context Learning (ICL). It conducts a detailed exploration of the types of shortcuts in ICL tasks, their causes, available benchmarks, and strategies for mitigating shortcuts. Based on corresponding observations, it summarizes the unresolved issues in existing research and attempts to outline the future research landscape of shortcut learning.

大模型捷径学习ICL泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。