arXiv:2410.16531cs.CLcs.AI2024-10被引 18

用贝叶斯视角解释上下文学习,揭示模型行为复现规律。

Bayesian scaling laws for in-context learning

  • 将上下文学习视为贝叶斯推断,建立可解释的预测模型。
  • 在不同规模GPT-2上验证,准确预测任务表现与参数关系。
  • 揭示后训练安全对齐失效原因,适合安全研究者参考。

上下文学习(ICL)是一种无需训练更新即可使语言模型完成复杂任务的强大技术。先前研究已发现提供示例数量与模型预测准确率之间存在强相关性。本文通过证明ICL近似于贝叶斯学习,提出一种新的贝叶斯尺度定律。在不同规模的GPT-2模型实验中,该定律不仅匹配现有准确率尺度规律,还为任务先验、学习效率和单样本概率提供了可解释项。为展示其分析能力,我们设计了受控合成数据集实验,用于指导真实世界中的安全对齐研究。实验中,先使用SFT或DPO抑制模型原有能力,再通过ICL尝试恢复该能力(多示例越狱)。随后在真实指令微调的大模型上,结合能力基准测试与新构建的多示例越狱数据集进行评估。结果表明,贝叶斯尺度定律能准确预测被抑制行为重新出现的条件,揭示了后训练提升大模型安全性的局限性。

原文摘要 · Abstract (English)

In-context learning (ICL) is a powerful technique for getting language models to perform complex tasks with no training updates. Prior work has established strong correlations between the number of in-context examples provided and the accuracy of the model's predictions. In this paper, we seek to explain this correlation by showing that ICL approximates a Bayesian learner. This perspective gives rise to a novel Bayesian scaling law for ICL. In experiments with \mbox{GPT-2} models of different sizes, our scaling law matches existing scaling laws in accuracy while also offering interpretable terms for task priors, learning efficiency, and per-example probabilities. To illustrate the analytic power that such interpretable scaling laws provide, we report on controlled synthetic dataset experiments designed to inform real-world studies of safety alignment. In our experimental protocol, we use SFT or DPO to suppress an unwanted existing model capability and then use ICL to try to bring that capability back (many-shot jailbreaking). We then study ICL on real-world instruction-tuned LLMs using capabilities benchmarks as well as a new many-shot jailbreaking dataset. In all cases, Bayesian scaling laws accurately predict the conditions under which ICL will cause suppressed behaviors to reemerge, which sheds light on the ineffectiveness of post-training at increasing LLM safety.

上下文学习贝叶斯推理模型安全尺度定律

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。