arXiv:2506.03292cs.CLcs.AI2025-06被引 20

用超网络实现大规模精准文本生成控制,无需为每条指令重训练。

HyperSteer: Activation Steering at Scale with Hypernetworks

  • 基于超网络动态生成控制向量,输入自然语言指令即刻响应。
  • 数千条指令下性能超越现有激活调控方法,包括未见过的指令。
  • 适合需要灵活、低成本控制生成内容的研究者和开发者使用。

通过修改语言模型内部激活值来控制文本生成是一种流行方法。无监督字典学习方法(如稀疏自编码器)可扩展生成大量控制向量,但无法保证每个向量的有效性,也难以覆盖相关控制任务。而有监督方法虽针对性强、效果好,却需为每个新增控制向量收集更多数据并重新训练。本文提出HyperSteer,一种基于超网络的端到端架构,能根据自然语言控制指令和被控语言模型内部状态,动态生成控制向量。评估显示,当使用数千个控制指令时,HyperSteer在性能上超过当前最先进的激活调控方法,甚至在训练中未见过的指令上也表现优异。此外,其性能与提示式控制相当。

原文摘要 · Abstract (English)

Steering language models (LMs) by modifying internal activations is a popular approach for controlling text generation. Unsupervised dictionary learning methods, e.g., sparse autoencoders, can be scaled to produce many steering vectors, but lack guarantees on the individual efficacy of each vector and control over the coverage of relevant steering tasks. In contrast, supervised methods for constructing steering vectors are targeted and effective, but require more data collection and training for each additional steering vector produced. In this work, we introduce HyperSteer, a family of hypernetwork-based architectures which are trained end-to-end to generate steering vectors conditioned on the natural language steering prompts and the internals of the steered LM. In our evaluations, we show that scaling HyperSteer with thousands of steering prompts exceeds the performance of state-of-the-art activation steering methods, even on steering prompts never seen during training. Moreover, HyperSteer performs on par with steering-via-prompting.

语言模型控制超网络生成调控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。