提出新指标LCS,精准评估大模型上下文学习效果。
Learning-to-Context Slope: Evaluating In-Context Learning Effectiveness Beyond Performance Illusions
- 用学习增益与上下文相关性的斜率衡量ICL效果。
- 在数据少或有偏时仍可靠,比性能变化更可信。
- 适合研究者和工程师优化模型上下文学习能力。
上下文学习(ICL)已成为提升大语言模型性能的有效方法。然而,其效果在不同模型和任务间差异显著,使从业者难以判断ICL是否真正有效。现有基于性能变化的评估方法存在可靠性低、归因性差、数据不足场景下不适用等问题。本文提出学习到上下文的斜率(LCS),通过建模学习增益(演示带来的损失下降)与上下文相关性(演示与输入的相关性)之间的斜率来量化ICL有效性。LCS克服了性能指标的三大局限:(1) 能捕捉输出错误时的连续损失变化,提升可靠性;(2) 可将ICL失败归因于弱上下文对齐(无法将输入适配演示)或强输出校准(自我验证正确性);(3) 通过合成评估减少对标注数据的依赖。大量实验表明,LCS在有标签场景中与性能提升强相关,在有偏差或数据稀缺场景中仍能真实反映有效性。进一步分析揭示了有效的LCS阈值,并识别出影响ICL成功的关键模型能力。
原文摘要 · Abstract (English)
In-context learning (ICL) has emerged as an effective approach to enhance the performance of large language models (LLMs). However, its effectiveness varies significantly across models and tasks, posing challenges for practitioners to determine when ICL reliably improves performance. Current evaluation approaches, reliant on performance change after applying ICL, suffer from low reliability, poor attribution, and impracticality in data-insufficient scenarios. We propose the Learning-to-Context Slope (LCS), a novel metric that quantifies ICL effectiveness by modeling the slope between learning gain (loss decrease from demonstrations) and contextual relevance (demonstration-input relevance). LCS addresses key limitations of performance-based metrics: (1) it captures continuous loss changes even when outputs are incorrect, improving reliability; (2) its formulation attributes ICL failures to weak contextual alignment (inability to adapt inputs to demonstrations) or strong output calibration (self-verification of correctness); and (3) it minimizes reliance on labeled data via synthetic evaluation. Extensive experiments demonstrate that LCS strongly correlates with performance improvements in labeled settings and reliably reflects true effectiveness in biased or data-scarce scenarios. Further analysis reveals actionable thresholds for LCS and identifies model capabilities critical to ICL success.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。