研究大模型中'诚实性'表示的可塑性演变,发现早期干预最有效。
A Timeline and Analysis for Representation Plasticity in Large Language Models
- 在不同微调阶段注入引导向量,观察模型行为变化
- 早期阶段可塑性强,后期存在响应敏感窗口
- 该规律跨模型架构通用,适合安全对齐干预
防止人工智能产生长期危险行为至关重要。表示工程(RepE)作为一种新兴方法,可在高层面上引导模型行为(如‘诚实性’)。理解此类表示的可塑性应成为对齐研究的核心。然而,当前对这一层面可塑性的研究严重不足。本文通过在不同微调阶段应用提取的引导向量,揭示了‘诚实性’表示稳定性与模型可塑性随时间的演化规律,发现早期阶段具有高可塑性,而后期存在出人意料的响应敏感窗口。该现象在多种模型架构中均被观测到,表明存在普遍的可塑性模式,可用于高效干预。研究成果极大促进人工智能透明性,解决当前制约行为引导效率的关键瓶颈。
原文摘要 · Abstract (English)
The ability to steer AI behavior is crucial to preventing its long term dangerous and catastrophic potential. Representation Engineering (RepE) has emerged as a novel, powerful method to steer internal model behaviors, such as "honesty", at a top-down level. Understanding the steering of representations should thus be placed at the forefront of alignment initiatives. Unfortunately, current efforts to understand plasticity at this level are highly neglected. This paper aims to bridge the knowledge gap and understand how LLM representation stability, specifically for the concept of "honesty", and model plasticity evolve by applying steering vectors extracted at different fine-tuning stages, revealing differing magnitudes of shifts in model behavior. The findings are pivotal, showing that while early steering exhibits high plasticity, later stages have a surprisingly responsive critical window. This pattern is observed across different model architectures, signaling that there is a general pattern of model plasticity that can be used for effective intervention. These insights greatly contribute to the field of AI transparency, addressing a pressing lack of efficiency limiting our ability to effectively steer model behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。