用微分方程统一建模大模型激活调控,实现更精准的对齐优化。
ODESteer: A Unified ODE-Based Steering Framework for LLM Alignment
- 基于常微分方程构建统一理论框架,将激活调控视为路径优化问题。
- 在TruthfulQA等三组基准上分别提升5.7%、2.5%、2.4%,效果稳定领先。
- 适合关注大模型对齐、可解释性与推理优化的研究者和实践者。
激活调控(或表征工程)通过在推理阶段操控大语言模型的内部激活,提供一种轻量级的对齐方法。然而,现有方法存在两大局限:(i) 缺乏统一的理论框架指导调控方向的设计;(ii) 过度依赖单步调控,难以捕捉激活分布的复杂模式。本文提出一种基于常微分方程(ODE)的统一理论框架,用于大模型对齐中的激活调控。我们证明,传统的激活加法可被视作该ODE解的一阶近似。基于此视角,确定调控方向等价于控制理论中的屏障函数设计。由此导出的ODESteer方法,通过将屏障函数定义为正负激活的对数密度比,构建多步自适应的调控路径。在多个大模型对齐基准测试中,相比现有先进方法,ODESteer表现出一致的性能提升:在TruthfulQA上提升5.7%,在UltraFeedback上提升2.5%,在RealToxicityPrompts上提升2.4%。本工作通过ODE统一了激活调控的理论基础,并以实证方式验证了其有效性。
原文摘要 · Abstract (English)
Activation steering, or representation engineering, offers a lightweight approach to align large language models (LLMs) by manipulating their internal activations at inference time. However, current methods suffer from two key limitations: (i) the lack of a unified theoretical framework for guiding the design of steering directions, and (ii) an over-reliance on one-step steering that fail to capture complex patterns of activation distributions. In this work, we propose a unified ordinary differential equations (ODEs)-based theoretical framework for activation steering in LLM alignment. We show that conventional activation addition can be interpreted as a first-order approximation to the solution of an ODE. Based on this ODE perspective, identifying a steering direction becomes equivalent to designing a barrier function from control theory. Derived from this framework, we introduce ODESteer, a kind of ODE-based steering guided by barrier functions, which shows empirical advancement in LLM alignment. ODESteer identifies steering directions by defining the barrier function as the log-density ratio between positive and negative activations, and employs it to construct an ODE for multi-step and adaptive steering. Compared to state-of-the-art activation steering methods, ODESteer achieves consistent empirical improvements on diverse LLM alignment benchmarks, a notable $5.7\%$ improvement over TruthfulQA, $2.5\%$ over UltraFeedback, and $2.4\%$ over RealToxicityPrompts. Our work establishes a principled new view of activation steering in LLM alignment by unifying its theoretical foundations via ODEs, and validating it empirically through the proposed ODESteer method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。