arXiv:2609.06879cs.CLcs.AI2026-09

自动构建词汇激活引导向量,精准控制大模型输出

AutoLexSteer: Automatic Contrast Construction for Lexical Activation Steering

论文配图:AutoLexSteer: Automatic Contrast Construction for Lexical Activation Steering
图 1 · 摘自论文原文
  • 基于词网同义词簇自动构造引导向量
  • 可精确控制词语及词义层面的模型行为
  • 适用于对抗模型谄媚等特定行为调控

引导向量已成为精准控制大语言模型输出的高效方法,但因嵌入空间不透明,手动构建准确向量仍具挑战。本文提出Hangman型引导向量,以词义为操作单元,并开发首个完全自动化构建流程AutoLexSteer。该方法利用WordNet提取紧密相关的词簇,分别指定需规避的源语义和期望的目标语义。生成的引导向量精度高,支持在词语及词义层级进行操控,可有效引导如谄媚等特定模型行为。代码与数据集见https://github.com/ShuheWang1998/autolexsteer。

原文摘要 · Abstract (English)

Steering vectors have rapidly emerged as a popular and effective method for guiding the output of LLMs in very specific ways. But constructing accurate steering vectors is a difficult manual process due to the opacity of embeddings. We introduce Hangman, a novel type of steering vector that operates using word senses, as well as AutoLexSteer, the first fully automated process for building steering vectors. AutoLexSteer employs families of closely-related words extracted from WordNet to specify both the steering source to be avoided and the desired steering target. The steering vectors are quite precise, can be used to steer at the level of words and sets of word senses (meanings), and are able to steer certain LLM behaviors like sycophancy. The dataset and code can be found at https://github.com/ShuheWang1998/autolexsteer.

大模型控制引导向量词义操控

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。