arXiv:2608.30986cs.LGcs.CL2026-08

通过几何优化实现大模型拒绝行为的高效控制

Controlling Refusal Behavior of LLMs via Stiefel-Constrained Rotation Steering

论文配图:Controlling Refusal Behavior of LLMs via Stiefel-Constrained Rotation Steering
图 1 · 摘自论文原文
  • 基于黎曼优化学习参数高效的旋转变换
  • 无需额外拒绝向量,干预效率显著提升
  • 适合需要精准控制大模型行为的研究者

激活值调制已成为推理时控制大模型拒绝行为的一种轻量级方法。越来越多的研究探索通过可训练的激活旋转来构建几何上合理的干预机制。然而,现有技术依赖于拒绝向量等辅助构造来定义旋转。本文提出一种自洽的方法,基于黎曼优化学习参数高效的旋转变换。实验验证了该方案在干预效率上的优越性。大量消融研究凸显了方法中关键设计选择的重要性。结果表明,所提出的基于旋转的调制方案是更可靠控制大模型行为的有前景方向。

原文摘要 · Abstract (English)

Activation steering has emerged as a lightweight approach for controlling model refusal at inference time. A growing line of research explores trainable rotations of activations to develop geometrically principled intervention mechanisms. However, existing techniques rely on auxiliary constructs, such as refusal vectors, to define these rotations. In our work, we develop a self-contained methodology for learning parameter-efficient rotational transformations based on Riemannian optimization. We empirically validate the proposed scheme, demonstrating its superiority in intervention efficiency. An extensive ablation study highlights the importance of key design choices in our method. Our results identify the proposed rotation-based steering scheme as a promising direction for more reliable control over the behavior of LLMs.

大模型控制激活调制几何优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。