发现语言模型在训练中段开始可被线性控制,且不同概念出现时间不同。
How Does Controllability Emerge In Language Models During Pretraining?
- 用统一框架分析隐藏状态,追踪控制能力随训练的演变
- 控制力在训练中段出现,且愤怒与悲伤等概念出现时机不同
- 提出热力图、熵趋势等指标,帮助理解控制力生成机制
语言模型可通过修改内部表示来控制生成内容的情感、风格或真实性。然而,有效干预的条件仍不明确,常依赖启发式方法和试错。本文表明,线性可控性(即通过线性变换隐藏状态调节输出的能力)在训练中期逐渐显现,且语义相近的概念(如愤怒与悲伤)在不同阶段才具备可控性。为此,我们提出“干预探测器”(Intervention Detector, ID),整合现有干预技术,通过隐藏状态与表征分析揭示线性可控性的动态演化。结果表明,随着训练推进,概念在隐藏空间中逐渐线性可分,这与可控性出现强相关。我们进一步引入热力图、熵趋势、余弦相似度等指标辅助解读可控性演化过程。该方法在多个模型族中验证,确保结论普适性。
原文摘要 · Abstract (English)
Language models can be steered by modifying their internal representations to control concepts such as emotion, style, or truthfulness in generation. However, the conditions for an effective intervention remain unclear and are often validated through heuristics and trial-and-error. To fill this gap, we demonstrate that intervention efficacy, measured by linear steerability (i.e., the ability to adjust output via linear transformations of hidden states), emerges during intermediate stages of training. Moreover, even closely related concepts (e.g., anger and sadness) exhibit steerability emergence at distinct stages of training. To better interpret the dynamics of steerability during training, we adapt existing intervention techniques into a unified framework, referred to as the "Intervention Detector" (ID), which is designed to reveal how linear steerability evolves over the course of training through hidden state and representation analysis. ID reveals that concepts become increasingly linearly separable in the hidden space as training progresses, which strongly correlates with the emergence of linear steerability. We further introduce ID-based metrics, such as heatmaps, entropy trends, and cosine similarity, to help interpret how linear steerability evolves throughout training. In addition, we apply ID across different model families to ensure the generality of our findings on steerability dynamics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。