用新方法追踪大模型如何从死记硬背变成理解语义
TRACE for Tracking the Emergence of Semantic Representations in Transformers
- 设计合成语料库和多维度指标,监测模型训练中抽象能力的演变
- 发现几何变化与语法语义准确率提升同步出现,标志关键转折点
- 结果在不同模型结构中稳定存在,适合研究模型可解释性与训练机制
现代Transformer模型在训练中会出现相变现象,即从记忆到抽象的明显转变,但其内在机制尚不清晰。以往研究多关注终点表示或孤立信号(如曲率、互信息),且局限于符号或算术任务,忽略了语言结构的涌现。本文提出TRACE(Tracking Representation Abstraction and Compositional Emergence)诊断框架,结合几何、信息与语言信号,检测基于Transformer的语言模型中的相变。该框架使用名为ABSynth的帧语义数据生成方法,构建可控复杂度、词频分布与结构熵的标注合成语料库,实现对抽象涌现的精准分析。实验表明:(i) 相变对应曲率坍缩与维度稳定化的明确交叉;(ii) 几何变化与句法、语义准确率提升同步;(iii) 抽象模式在不同架构中持续存在,前馈网络影响优化稳定性而非根本改变演化轨迹。本研究深化了对语言抽象如何在模型中涌现的理解,为模型可解释性、训练效率与组合泛化提供洞见,有助于更系统地设计语言模型。
原文摘要 · Abstract (English)
Modern transformer models exhibit phase transitions during training, distinct shifts from memorisation to abstraction, but the mechanisms underlying these transitions remain poorly understood. Prior work has often focused on endpoint representations or isolated signals like curvature or mutual information, typically in symbolic or arithmetic domains, overlooking the emergence of linguistic structure. We introduce TRACE (Tracking Representation Abstraction and Compositional Emergence), a diagnostic framework combining geometric, informational, and linguistic signals to detect phase transitions in Transformer-based LMs. TRACE leverages a frame-semantic data generation method, ABSynth, that produces annotated synthetic corpora with controllable complexity, lexical distributions, and structural entropy, while being fully annotated with linguistic categories, enabling precise analysis of abstraction emergence. Experiments reveal that (i) phase transitions align with clear intersections between curvature collapse and dimension stabilisation; (ii) these geometric shifts coincide with emerging syntactic and semantic accuracy; (iii) abstraction patterns persist across architectural variants, with components like feedforward networks affecting optimisation stability rather than fundamentally altering trajectories. This work advances our understanding of how linguistic abstractions emerge in LMs, offering insights into model interpretability, training efficiency, and compositional generalisation that could inform more principled approaches to LM development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。