用内在维度衡量大模型中的语言复杂度,发现不同句法结构对应不同复杂度特征。
Tracing the complexity profiles of different linguistic phenomena through the intrinsic dimension of LLM representations
- 通过内在维度分析模型各层表征,捕捉语言复杂度差异。
- 复杂句法结构(如中心嵌套)在模型中引发更高内在维度。
- 不同复杂度类型在模型不同层级达到峰值,提示共通处理机制。
我们探索大语言模型表征的内在维度(ID)作为语言复杂度的指标。具体而言,测试了模型层间ID差异是否反映心理语言学中已确立的语言复杂度对比:并列与从属、右分支与中心嵌套、明确与模糊依附。在六种不同大语言模型上的实验结果表明,这些复杂度对比在ID差异中得到一致体现,更复杂的语言现象诱发更高的内在维度。值得注意的是,不同对比的ID差异在不同层段显现,并在不同阶段达到峰值。使用表征相似性和层剪枝的进一步实验验证了这一趋势。结论表明,内在维度是大模型中语言复杂度的有用指标,能揭示跨不同模型的相似语言处理阶段,并具备区分不同类型复杂度的潜力。
原文摘要 · Abstract (English)
We explore intrinsic dimension (ID) of LLM representations as a marker of linguistic complexity. Specifically, we test whether ID differences across model layers reflect well-known complexity contrasts established in (psycho)linguistics: coordination vs. subordination, right-branching vs. center-embedding, and unambiguous vs. ambiguous attachment. Our results on six different LLMs show that these contrasts are consistently reflected in ID differences, with more complex phenomena eliciting higher ID profiles. Notably, ID differences emerge at different points across layers for different contrasts, also reaching their peaks at different stages. Further experiments using representational similarity and layer pruning confirm the trends. We conclude that ID is a useful marker of linguistic complexity in LLMs, that it points to similar linguistic processing steps across disparate LLMs, and that it has the potential to differentiate between different types of complexity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。