揭示语言模型如何从词元自发形成概念格结构。
From Tokens to Lattices: Emergent Lattice Structures in Language Models
- 用形式概念分析法挖掘模型隐含的概念关系。
- 实证发现预训练模型能自动构建概念格结构。
- 无需人工定义概念,可发现潜在新概念。
预训练掩码语言模型(MLMs)展现出强大的概念知识理解与编码能力,揭示出概念间存在格结构。本文从形式概念分析(FCA)视角探讨这一概念化如何从预训练中产生。我们证明,MLM的训练目标隐式学习了一个描述对象、属性及其依赖关系的“形式上下文”,使通过FCA重构概念格成为可能。提出一种基于预训练MLM的新概念格构建框架,不依赖人类定义的概念,可发现超越人类认知的“潜在概念”。构建三个数据集进行评估,实验结果验证了该假设。
原文摘要 · Abstract (English)
Pretrained masked language models (MLMs) have demonstrated an impressive capability to comprehend and encode conceptual knowledge, revealing a lattice structure among concepts. This raises a critical question: how does this conceptualization emerge from MLM pretraining? In this paper, we explore this problem from the perspective of Formal Concept Analysis (FCA), a mathematical framework that derives concept lattices from the observations of object-attribute relationships. We show that the MLM's objective implicitly learns a \emph{formal context} that describes objects, attributes, and their dependencies, which enables the reconstruction of a concept lattice through FCA. We propose a novel framework for concept lattice construction from pretrained MLMs and investigate the origin of the inductive biases of MLMs in lattice structure learning. Our framework differs from previous work because it does not rely on human-defined concepts and allows for discovering "latent" concepts that extend beyond human definitions. We create three datasets for evaluation, and the empirical results verify our hypothesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。