arXiv:2508.08673stat.MLcs.LG2025-08被引 1

揭示了上下文学习的最优概率估计原理,证明其在多分类中达到理论极限。

In-Context Learning as Nonparametric Conditional Probability Estimation: Risk Bounds and Optimality

  • 将上下文学习视为条件概率估计,用KL散度衡量风险。
  • 证明Transformer和MLP在特定条件下都能达到最优估计率(对数因子内)。
  • 提出新方法控制泛化误差,适用于多种模型架构。

本文研究多分类任务中上下文学习(ICL)的期望超额风险。将每个任务形式化为一系列带标签样本后接一个查询输入;预训练模型据此估计查询的条件类别概率分布。期望超额风险定义为在指定任务族上,预测与真实条件类别分布之间的平均截断KL散度。本文建立了一个基于KL散度的新奥拉克不等式,适用于多分类场景。该结果导出了Transformer模型的紧致上下界,表明ICL估计器在条件概率估计中达到最小最大最优率(对数因子内)。技术上,本工作引入一种通过统一经验熵控制泛化误差的新方法。进一步证明,在合适假设下,多层感知机(MLPs)也能实现ICL并达到相同最优率,表明有效上下文学习不必仅限于Transformer架构。

原文摘要 · Abstract (English)

This paper investigates the expected excess risk of in-context learning (ICL) for multiclass classification. We formalize each task as a sequence of labeled examples followed by a query input; a pretrained model then estimates the query's conditional class probabilities. The expected excess risk is defined as the average truncated Kullback-Leibler (KL) divergence between the predicted and true conditional class distributions over a specified family of tasks. We establish a new oracle inequality for this risk, based on KL divergence, in multiclass classification. This yields tight upper and lower bounds for transformer-based models, showing that the ICL estimator achieves the minimax optimal rate (up to logarithmic factors) for conditional probability estimation. From a technical standpoint, our results introduce a novel method for controlling generalization error via uniform empirical entropy. We further demonstrate that multilayer perceptrons (MLPs) can also perform ICL and attain the same optimal rate (up to logarithmic factors) under suitable assumptions, suggesting that effective ICL need not be exclusive to transformer architectures.

上下文学习概率估计最优率Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。