研究大模型如何识别实体,发现其内部表示能有效区分不同实体。
On Entity Identification in Language Models
- 通过聚类分析语言模型内部表征,量化实体提及的聚集与分离程度。
- 五种模型在实体识别上表现良好,准确率与召回率均达0.66至0.9之间。
- 早期层中实体信息集中在低维线性子空间,适合研究模型知识结构。
我们分析了语言模型(LMs)内部表征在识别和区分命名实体提及方面的能力,重点关注实体与其提及之间的多对多关系。首先提出实体提及的两个问题——歧义性和变异性,并构建类聚类质量评估框架。通过分析模型内部表征的聚类效果,量化同一实体的提及是否聚合、不同实体的提及是否分离。实验涵盖五种基于Transformer的自回归模型,结果显示其在实体识别上的精度与召回率分别达到0.66至0.9。进一步分析表明,实体相关信息在早期层中以低维线性子空间形式紧凑表示。此外,揭示了实体表征特性对词预测性能的影响。这些发现从语言模型表征与真实世界实体知识结构同构性的角度进行解释,为理解模型如何组织和使用实体信息提供了新视角。
原文摘要 · Abstract (English)
We analyze the extent to which internal representations of language models (LMs) identify and distinguish mentions of named entities, focusing on the many-to-many correspondence between entities and their mentions. We first formulate two problems of entity mentions -- ambiguity and variability -- and propose a framework analogous to clustering quality metrics. Specifically, we quantify through cluster analysis of LM internal representations the extent to which mentions of the same entity cluster together and mentions of different entities remain separated. Our experiments examine five Transformer-based autoregressive models, showing that they effectively identify and distinguish entities with metrics analogous to precision and recall ranging from 0.66 to 0.9. Further analysis reveals that entity-related information is compactly represented in a low-dimensional linear subspace at early LM layers. Additionally, we clarify how the characteristics of entity representations influence word prediction performance. These findings are interpreted through the lens of isomorphism between LM representations and entity-centric knowledge structures in the real world, providing insights into how LMs internally organize and use entity information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。