揭示大模型内部语义角色电路的形成与定位规律
Emergence and Localisation of Semantic Role Circuits in LLMs
- 通过最小对比、时间演化和跨模型对比,追踪语义角色实现路径
- 89%-94%的语义贡献集中在28个节点内,结构高度集中
- 电路在不同模型间部分共享,但大型模型可能绕过局部电路
尽管大型语言模型展现出语义能力,但其内部如何构建抽象语义结构仍不明确。本文提出一种结合角色交叉最小对、时间演化分析和跨模型比较的方法,研究大模型如何实现语义角色。分析发现:(i) 语义贡献高度集中于少数节点(89%-94%归因于28个节点);(ii) 结构呈现渐进式优化,而非突变式跃迁,且更大模型有时会跳过局部电路;(iii) 不同尺度间存在中等程度的组件重叠(24%-59%),但频谱相似性高。结果表明,大模型以紧凑、因果隔离的方式构建抽象语义机制,且该机制在不同规模与架构间具有部分可迁移性。
原文摘要 · Abstract (English)
Despite displaying semantic competence, large language models' internal mechanisms that ground abstract semantic structure remain insufficiently characterised. We propose a method integrating role-cross minimal pairs, temporal emergence analysis, and cross-model comparison to study how LLMs implement semantic roles. Our analysis uncovers: (i) highly concentrated circuits (89-94% attribution within 28 nodes); (ii) gradual structural refinement rather than phase transitions, with larger models sometimes bypassing localised circuits; and (iii) moderate cross-scale conservation (24-59% component overlap) alongside high spectral similarity. These findings suggest that LLMs form compact, causally isolated mechanisms for abstract semantic structure, and these mechanisms exhibit partial transfer across scales and architectures.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。