探究大模型如何组合子词信息,揭示其内部表征机制差异
Understanding Subword Compositionality of Large Language Models
- 通过三类实验分析子词组合的结构、语义与形式特征
- 发现五类模型在层间演化中呈现三种不同组合模式
- 适合关注模型内部工作机制的研究者参考
大型语言模型(LLMs)以子词序列为输入,需有效将子词表示组合为有意义的词级表示。本文通过一系列实验,系统探查了LLMs在结构相似性、语义可分解性和形式保留性三个关键方面的子词组合能力。对五类主流模型家族的分析表明,它们可被划分为三类不同组别,反映出底层组合策略的差异。具体发现:(i)子词组合与完整词表示之间的结构相似性在各层演化中存在三种明显模式;(ii)逐层探测显示模型对语义可分解性的敏感度极高;(iii)对形式特征(如字符序列长度)的敏感性也表现出三种不同模式。这些结果揭示了大模型在子词信息编码与整合中的动态机制,凸显了不同模型在组合方式上的多样性。
原文摘要 · Abstract (English)
Large language models (LLMs) take sequences of subwords as input, requiring them to effective compose subword representations into meaningful word-level representations. In this paper, we present a comprehensive set of experiments to probe how LLMs compose subword information, focusing on three key aspects: structural similarity, semantic decomposability, and form retention. Our analysis of the experiments suggests that these five LLM families can be classified into three distinct groups, likely reflecting difference in their underlying composition strategies. Specifically, we observe (i) three distinct patterns in the evolution of structural similarity between subword compositions and whole-word representations across layers; (ii) great performance when probing layer by layer their sensitivity to semantic decompositionality; and (iii) three distinct patterns when probing sensitivity to formal features, e.g., character sequence length. These findings provide valuable insights into the compositional dynamics of LLMs and highlight different compositional pattens in how LLMs encode and integrate subword information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。