arXiv:2411.07130cs.CL2024-11ACL被引 11

提出新评估框架,测试大模型长文本理解能力。

On Many-Shot In-Context Learning for Long-Context Evaluation

  • 区分任务类型:相似样本学习与全样本理解
  • 64k tokens内分类/摘要任务表现好,16k token时推理任务骤降
  • 适合研究长上下文建模与评估的学者

多示例上下文学习(Many-shot ICL)已成为评估大语言模型处理长上下文能力的独特方式。本文通过多示例ICL深入研究长上下文语言模型(LCLM)的评估方法。我们发现,分类与摘要任务在增加示范样本后性能提升,而翻译和推理任务则无明显趋势。进一步分析表明,不同任务对检索相似样本或全局上下文理解的需求不同,据此提出两类任务划分:(i) 相似样本学习(SSL),仅需检索最相似示例即可;(ii) 全样本学习(ASL),需全面理解提示中所有示例。为此构建新基准MANYICLBENCH,评测12个LCLM。结果显示,先进模型在SSL任务中可稳定支持64k tokens,但在ASL任务中,多数模型在16k tokens时即出现显著性能下降。

原文摘要 · Abstract (English)

Many-shot in-context learning (ICL) has emerged as a unique setup to both utilize and test the ability of large language models to handle long context. This paper delves into long-context language model (LCLM) evaluation through many-shot ICL. We first ask: what types of ICL tasks benefit from additional demonstrations, and how effective are they in evaluating LCLMs? We find that classification and summarization tasks show performance improvements with additional demonstrations, while translation and reasoning tasks do not exhibit clear trends. Next, we investigate the extent to which different tasks necessitate retrieval versus global context understanding. We develop metrics to categorize ICL tasks into two groups: (i) similar-sample learning (SSL): tasks where retrieval of the most similar examples is sufficient for good performance, and (ii) all-sample learning (ASL): tasks that necessitate a deeper comprehension of all examples in the prompt. Lastly, we introduce a new many-shot ICL benchmark, MANYICLBENCH, to characterize model's ability on both fronts and benchmark 12 LCLMs using MANYICLBENCH. We find that while state-of-the-art models demonstrate good performance up to 64k tokens in SSL tasks, many models experience significant performance drops at only 16k tokens in ASL tasks.

长上下文评估基准ICL模型测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。