arXiv:2603.12298cs.LGcs.AI2026-03被引 5

通过全局演化稳定性优化激活控制,提升大模型指令对齐可靠性。

Global Evolutionary Steering: Refining Activation Steering Control via Cross-Layer Consistency

  • 基于网络表征演化的几何稳定性,修正原始激活向量
  • 在多个任务上超越基线,无需分层调参即保持稳定表现
  • 适合需要无训练微调的精准控制场景

激活工程可在不进行微调的情况下实现对大语言模型的精确控制。然而,现有方法从静态激活差异中提取向量,易受高维噪声和层间语义漂移影响,常捕捉到虚假相关而非目标意图。为此,我们提出训练无关的全局演化精炼控制(GER-steer),其基于网络表征演化的几何稳定性。GER-steer利用这一全局信号纠正原始控制向量,有效分离出稳健的语义意图与正交干扰。大量实验表明,GER-steer持续优于基线,在无需分层调参的情况下实现更优效能与泛化能力,为可靠模型对齐提供通用解决方案。

原文摘要 · Abstract (English)

Activation engineering enables precise control over Large Language Models (LLMs) without the computational cost of fine-tuning. However, existing methods deriving vectors from static activation differences are susceptible to high-dimensional noise and layer-wise semantic drift, often capturing spurious correlations rather than the target intent. To address this, we propose Global Evolutionary Refined Steering (GER-steer), a training-free framework that grounded in the geometric stability of the network's representation evolution. GER-steer exploits this global signal to rectify raw steering vectors, effectively decoupling robust semantic intent from orthogonal artifacts. Extensive evaluations confirm that GER-steer consistently outperforms baselines, delivering superior efficacy and generalization without layer-specific tuning, establishing a universal solution for reliable model alignment.

激活工程模型对齐无训练微调

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。