arXiv:2410.01769cs.CL2024-10被引 21

提出新评估框架,量化大模型在复杂任务中的泛化能力边界。

Quantifying Generalization Complexity for Large Language Models

  • 通过20个任务、5级复杂度,区分模型在分布内/外的表现差异。
  • 发现复杂度与泛化差距呈非单调关系,存在关键阈值点。
  • 适用于评估开源与闭源大模型的泛化极限,指导模型设计。

尽管大语言模型(LLMs)在理解复杂问题和执行高阶任务方面表现出色,但其泛化能力常与记忆行为纠缠,亟需更精确的评估方法。为此,我们提出Scylla——一个动态评估框架,可定量测量LLMs的泛化能力。该框架通过20项任务、5个复杂度层级,分别评估模型在分布内(ID)与分布外(OOD)数据上的表现,有效分离泛化与记忆。实验揭示任务复杂度与ID/OOD性能差距之间存在非单调关系,称为‘泛化谷’现象。具体而言,在某一临界复杂度时,模型对非泛化行为的依赖达到峰值,标志着其泛化能力的上限。随着模型规模增大,该临界复杂度向更高复杂度移动,表明更大模型可在更复杂推理任务中维持泛化能力。基于Scylla和临界复杂度概念,我们对28个LLMs(包括LLaMA、Qwen等开源模型及Claude、GPT等闭源模型)进行了基准测试,提供了更可靠的评估体系,深化了对大模型泛化能力的理解。

原文摘要 · Abstract (English)

While large language models (LLMs) have shown exceptional capabilities in understanding complex queries and performing sophisticated tasks, their generalization abilities are often deeply entangled with memorization, necessitating more precise evaluation. To address this challenge, we introduce Scylla, a dynamic evaluation framework that quantitatively measures the generalization abilities of LLMs. Scylla disentangles generalization from memorization via assessing model performance on both in-distribution (ID) and out-of-distribution (OOD) data through 20 tasks across 5 levels of complexity. Through extensive experiments, we uncover a non-monotonic relationship between task complexity and the performance gap between ID and OOD data, which we term the generalization valley. Specifically, this phenomenon reveals a critical threshold - referred to as critical complexity - where reliance on non-generalizable behavior peaks, indicating the upper bound of LLMs' generalization capabilities. As model size increases, the critical complexity shifts toward higher levels of task complexity, suggesting that larger models can handle more complex reasoning tasks before over-relying on memorization. Leveraging Scylla and the concept of critical complexity, we benchmark 28LLMs including both open-sourced models such as LLaMA and Qwen families, and close-sourced models like Claude and GPT, providing a more robust evaluation and establishing a clearer understanding of LLMs' generalization capabilities.

大模型泛化能力评估框架复杂度

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。