用几何引导搜索,让大模型概念控制更高效省力。
When Is Rank-1 Steering Cheap? Geometry, Granularity, and Budgeted Search

- 基于激活几何设计预算约束的搜索策略,减少试错次数。
- 高概念粒度导致方向不一致,需更多尝试才能找到最优干预。
- 提出GRACE框架,自动诊断困难根源并分配优化资源。
激活操控为无需重训练即可控制大语言模型提供了轻量方法,但其效果在不同概念间差异显著。我们认为这种差异主要源于搜索难度:有效的一维干预方向虽存在,但寻找它可能代价高昂。本文将秩-1操控形式化为对干预层和系数的预算约束优化问题。实验表明,提示边界方向对齐能预测有效干预位置,实现几何引导搜索,在三个模型家族上平均减少39.8%的评估次数即可达到95%的最佳性能。进一步引入概念粒度(granularity)衡量对比上下文中的方向异质性,发现高粒度与慢收敛、低最佳性能相关(皮尔逊相关系数r=0.44, p<0.001;r=-0.46, p<0.001)。为此提出GRACE框架,利用激活几何诊断困难来源,选择合适修复方式,并高效分配优化资源。研究视角从‘何时秩-1失败’转向‘何时秩-1廉价且稳定’,使激活几何成为可行动的先验知识。
原文摘要 · Abstract (English)
Activation steering offers a lightweight way to control LLMs without retraining, but its effectiveness varies sharply across concepts. Prior work often reads this variability as evidence that many concepts are not captured by a single steering direction. We argue instead that much of it reflects search difficulty: a useful rank-1 intervention often exists, but finding it can be expensive. We formalize rank-1 steering as a budget-constrained optimization over intervention layer and coefficient. Across concepts and model families, prompt-boundary directional alignment predicts where effective interventions occur, enabling geometry-guided search that reaches high utility with substantially fewer evaluations, reducing the trials needed to recover 95% of best-found utility by 39.8% on average across three model families. To explain why some concepts remain expensive even under better search, we introduce concept granularity, a measure of directional heterogeneity across contrastive contexts. Granularity distinguishes concepts whose difference vectors share a stable global direction from those where prompts agree locally within each input but the utility-maximizing direction rotates systematically across inputs. Higher granularity is associated with slower convergence and lower best-found performance (Pearson $r{=}0.44$ with trials-to-95%, $r{=}{-}0.46$ with best-found utility, both $p<0.001$). We present GRACE, a Granularity- and Representation-Aware Concept Engineering framework that uses activation geometry to diagnose the dominant source of steering difficulty, select the appropriate remedy, and allocate optimization effort efficiently. Our results shift the frame from "when does rank-1 fail?" to "when is rank-1 cheap and stable?", turning activation geometry from a descriptive tool into an actionable prior for LLM control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。