arXiv:2409.13728cs.CLcs.LG2024-09中稿 · NeurIPS被引 2

研究大模型在陌生规则下的推理能力,揭示其泛化机制。

Rule Extrapolation in Language Models: A Study of Compositional Generalization on OOD Prompts

  • 用形式语言构建规则违反的分布外场景,测试模型推理能力。
  • 发现Transformer在复杂规则下表现优于线性与循环结构。
  • 提出基于算法信息论的理论框架,解释模型外推规律。

大型语言模型展现出惊人涌现能力,例如从看似分布外的提示中推断概念,即上下文学习。尽管这一成功常归因于Transformer架构,但我们的系统性理解仍有限。在复杂的现实数据集中,定义什么是分布外也并不明确。为更深入理解自回归大模型在分布外行为的表现,我们聚焦于由规则交集定义的形式语言。我们提出一种新的分布外组合泛化场景,称为规则外推。规则外推描述的是提示违反至少一条规则的分布外情形。我们在不同复杂度的线性、递归架构、Transformer及状态空间模型上评估了规则外推能力,以分析架构对规则外推的影响。同时,我们首次建立规则外推的规范理论基础,灵感源自算法信息论中的索洛莫诺夫先验。

原文摘要 · Abstract (English)

LLMs show remarkable emergent abilities, such as inferring concepts from presumably out-of-distribution prompts, known as in-context learning. Though this success is often attributed to the Transformer architecture, our systematic understanding is limited. In complex real-world data sets, even defining what is out-of-distribution is not obvious. To better understand the OOD behaviour of autoregressive LLMs, we focus on formal languages, which are defined by the intersection of rules. We define a new scenario of OOD compositional generalization, termed rule extrapolation. Rule extrapolation describes OOD scenarios, where the prompt violates at least one rule. We evaluate rule extrapolation in formal languages with varying complexity in linear and recurrent architectures, the Transformer, and state space models to understand the architectures' influence on rule extrapolation. We also lay the first stones of a normative theory of rule extrapolation, inspired by the Solomonoff prior in algorithmic information theory.

规则外推大模型泛化能力

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。