纯记忆模型也能伪装成先学抽象,颠覆大模型学习顺序的判断标准。
Exemplars in Disguise: Pure Exemplar Models Mimic Abstraction-First Learning

- 用纯记忆模型模拟学习过程,验证学习顺序判断方法的不可靠性。
- 模型学习顺序取决于输入分布和对个体样本的敏感度,而非真实抽象能力。
- 分布式表示下,具体与抽象知识难以区分,理论基础存疑。
个体特异性知识是先于还是后于类别级抽象知识被学习,是语言学习中的核心问题,示例模型与抽象理论对此有相反预测。近期研究声称大语言模型先习得抽象知识。本文指出这些方法存在缺陷:无抽象表征的纯记忆模型,依据相同评判标准,可能表现出先学具体或先学类别知识,取决于对个体观察的敏感度,而转折点由输入分布决定。进一步认为,在分布式表示中,词的类别属性与个体属性难以分离,导致具体与抽象知识的区分可能不成立。
原文摘要 · Abstract (English)
Whether idiosyncratic, item-specific knowledge is learned before abstract class-level generalizations, or vice versa, is a central question in language learning, with exemplar and abstraction-based theories making opposite predictions. Recent methods have claimed to show that, at least for large language models, abstract knowledge is learned first. We show that these methods fall short: pure memorizer models with no abstract representations can appear, by the same criteria, to learn either item-specific or class-level knowledge first, depending on their sensitivity to individual observations, with the transition point governed by the distributional properties of the input. We further argue that the distinction between item-specific and abstract knowledge may be ill-defined for distributed representations, as a word's class-level properties may not be separable from its item-specific properties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。