arXiv:2505.13089cs.CL2025-05ACL被引 3

用信息熵衡量语言模型系统泛化能力,发现熵越高模型表现越好

Systematic Generalization in Language Models Scales with Information Entropy

  • 用组件分布熵量化系统泛化的难度
  • 主流模型性能随训练数据熵升高而提升
  • 为评估模型泛化能力提供新指标,适合关注模型鲁棒性的研究者

当前语言模型在系统泛化方面仍面临挑战,对输入的语义相似变换敏感,且在新情境中难以处理已知概念。尽管已有基准评估组合行为,但尚无明确方法衡量系统泛化问题的难度。本文表明,系统泛化的一个方面可通过训练数据中组件部分分布的熵来描述。我们构建了序列到序列任务中的熵测量框架,发现主流模型架构的性能随熵值升高而提升。该工作将系统泛化与信息效率联系起来,结果表明:即使无内置先验,高熵场景下也能取得成功;而低熵场景的成功可作为评估鲁棒系统泛化进展的目标。

原文摘要 · Abstract (English)

Systematic generalization remains challenging for current language models, which are known to be both sensitive to semantically similar permutations of the input and to struggle with known concepts presented in novel contexts. Although benchmarks exist for assessing compositional behavior, it is unclear how to measure the difficulty of a systematic generalization problem. In this work, we show how one aspect of systematic generalization can be described by the entropy of the distribution of component parts in the training data. We formalize a framework for measuring entropy in a sequence-to-sequence task and find that the performance of popular model architectures scales with the entropy. Our work connects systematic generalization to information efficiency, and our results indicate that success at high entropy can be achieved even without built-in priors, and that success at low entropy can serve as a target for assessing progress towards robust systematic generalization.

语言模型系统泛化信息熵

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。