arXiv:2605.23040cs.LG2026-05

通过稀疏查询特征优化,实现大模型生成的精准可控。

Steered Generation via Gradient-Based Optimization on Sparse Query Features

论文配图:Steered Generation via Gradient-Based Optimization on Sparse Query Features
图 1 · 摘自论文原文
  • 用稀疏自编码器分解注意力查询,提取可解释特征
  • 在网格世界中成功引导安全与最短路径规划
  • 适用于逻辑推理与风格控制,适合需要精确调控的场景

隐空间操控利用大语言模型内部表示来引导生成,但对密集状态的干预容易导致语义特征纠缠。本文研究注意力查询激活作为高保真控制点,假设操纵注意力机制本身比通用状态干预更具精确性。提出基于原型的稀疏操控框架,将稀疏自编码器(SAEs)专门应用于查询激活,将其分解为可解释特征,并在推理阶段通过梯度优化使稀疏表示与目标行为的类别原型对齐。为验证该架构洞察,首先在文本化网格世界(Textualized Gridworld)这一可验证规划约束的受控环境中进行分析。结果表明,优化稀疏查询特征能有效满足刚性规划要求(如安全路径与最短路径),证实方法可实现客观规则约束。随后,通过在高维教育领域训练SAEs,展示框架可调控反馈的认知复杂度(即布卢姆分类学层级)。实验表明,稀疏查询表示提供了统一且可解释的控制能力,能同时处理逻辑规划与风格细微差别。

原文摘要 · Abstract (English)

Latent steering exploits internal representations of Large Language Models (LLMs) to guide generation, yet interventions on dense states can entangle distinct semantic features. In this paper, we investigate attention query activations as a high-fidelity site for precise control, hypothesizing that manipulating the attention mechanism itself offers sharper steerability than general state interventions. We introduce Prototype-Based Sparse Steering, a framework that applies Sparse Autoencoders (SAEs) specifically to query activations, to decompose them into interpretable features, then apply gradient-based optimization during inference to align the sparse representation with class prototypes of target behaviors. To validate this architectural insight, we first analyze the mechanism in Textualized Gridworld, a controlled environment for verifiable planning constraints. We demonstrate that optimizing sparse query features enables effective navigation of rigid planning requirements (i.e., safe vs. short paths), confirming the method's ability to satisfy objective rules. We then demonstrate the framework's versatility by training SAEs on a high-dimensional educational domain, where the framework steers the cognitive complexity of feedback (i.e., Bloom's Taxonomy). Our experiments establish that sparse query representations provide the necessary disentanglement for unified, interpretable control over both logical planning and stylistic nuance.

大模型控制注意力机制稀疏表示

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。