arXiv:2502.04580cs.LGcs.AI2025-02NeurIPS

发现长上下文下模型学习效率下降,揭示ICL的固有缺陷

Technical Debt in In-Context Learning: Diminishing Efficiency in Long Context

  • 构建元ICL框架,测试模型在不同任务中的学习效率
  • 长上下文时ICL效率显著低于贝叶斯最优解,样本需求翻倍
  • 适合关注大模型泛化能力与高效学习机制的研究者

Transformer在上下文学习(ICL)中表现出色,能通过示例直接适应新任务而无需参数更新。尽管已有理论和实证表明其可能优于特定任务模型,但其在上下文学习中的最优性仍不明确。为此,本文提出一种元ICL框架,每个提示定义一个回归任务,目标函数来自分层分布,需同时推断潜在模型类与任务特定参数。在此设定下,对比了ICL与贝叶斯最优估计器等原理性算法的样本复杂度,评估在不同性能要求下的表现。结果揭示显著二分:初期ICL效率接近贝叶斯最优,但在长上下文时效率大幅下降。信息论分析表明,此效率衰减是ICL固有特性。研究澄清了使用ICL作为通用求解器的权衡,推动开发无效率衰减的实时自适应方法。

原文摘要 · Abstract (English)

Transformers have demonstrated remarkable in-context learning (ICL) capabilities, adapting to new tasks by simply conditioning on demonstrations without parameter updates. Compelling empirical and theoretical evidence suggests that ICL, as a general-purpose learner, could outperform task-specific models. However, it remains unclear to what extent the transformers optimally learn in-context compared to principled learning algorithms. To investigate this, we employ a meta ICL framework in which each prompt defines a distinctive regression task whose target function is drawn from a hierarchical distribution, requiring inference over both the latent model class and task-specific parameters. Within this setup, we benchmark sample complexity of ICL against principled learning algorithms, including the Bayes optimal estimator, under varying performance requirements. Our findings reveal a striking dichotomy: while ICL initially matches the efficiency of a Bayes optimal estimator, its efficiency significantly deteriorates in long context. Through an information-theoretic analysis, we show that the diminishing efficiency is inherent to ICL. These results clarify the trade-offs in adopting ICL as a universal problem solver, motivating a new generation of on-the-fly adaptive methods without the diminishing efficiency.

上下文学习模型效率贝叶斯推理Transformer

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。