arXiv:2606.15092cs.LG2026-06

通过高维随机投影提升语言模型行为控制能力

High-Dimensional Random Projection for Activation Steering in Language Models

  • 在高维空间中进行激活值投影加法,捕捉非线性特征结构
  • 在多个大模型和基准测试中表现优于基线方法
  • 无需训练,可无缝集成现有控制技术,适合快速部署

激活操控已成为控制大语言模型行为的关键方法。然而,现有的基于均值差异的方法存在根本局限:仅能捕捉类别激活的均值差异,无法恢复超叠加假设下存在于非线性特征子空间中的判别信号。为此,我们提出一种无需训练的高维随机投影激活操控方法(HiDRA),可无缝集成现有激活操控技术。通过在投影的高维空间中执行激活加法,HiDRA 能够严格证明地捕获线性方法无法触及的更优判别结构。在多种大模型家族和基准测试上的实验表明,HiDRA 均显著优于基线方法,在不增加明显计算开销的前提下实现更强的行为控制能力。

原文摘要 · Abstract (English)

Activation steering has emerged as a key methodology for controlling the behavior of large language models (LLMs). Existing difference-in-means based methods, however, are fundamentally limited: they capture only mean differences between class activations and fail to recover discriminative signals that naturally exist in the nonlinear feature subspace under the superposition hypothesis. Motivated by that, we propose High-Dimensional Random-projection for Activation Steering (HiDRA), a training-free approach that integrates seamlessly with existing activation steering methods. By performing activation addition in the projected high-dimensional space, HiDRA can provably capture a better discriminative structure beyond the reach of linear methods. Experiments across diverse LLM families and benchmarks demonstrate that HiDRA consistently outperforms baseline counterparts, achieving stronger behavioral control without significant computational overhead.

激活操控大模型控制随机投影

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。