arXiv:2607.05790cs.AI2026-07被引 2

通过特定标题向量操控大模型工具使用,有效减少无关调用。

Controlling Tool Use with Heading-Specific Activation Steering

论文配图:Controlling Tool Use with Heading-Specific Activation Steering
图 1 · 摘自论文原文
  • 从标题锚点提取引导向量,控制模型是否调用外部工具。
  • 在五个开源模型上验证,可抑制无需工具的场景下90%以上的误调用。
  • 发现工具调用具非线性几何结构,适合研究工具使用机制的人参考。

具备工具增强能力的大语言模型可通过外部工具拓展其能力,但常出现不必要的工具调用。本文探究工具调用决策是否存在可提取、可操纵的稳定内部表征。尽管工具仅在推理时存在于上下文,不直接编码于模型权重中,我们发现从标题锚点位置提取的引导向量,可在五种开源模型及三个领域中对工具调用行为实现双向因果控制。在仅靠模型参数即可完成推理的领域,该方法抑制无效调用的效果最佳。然而几何分析显示,这种因果有效性并未对应清晰的线性结构:工具调用步骤与抑制向量呈现扩散、双峰对齐,而非线性理论预期的一致负相关;不同工具类型激活的内部特征差异大,跨工具特征重叠低。我们推测这些几何特性反映了工具的非参数本质,且将工具使用引导向量与参数化概念引导向量区分开来。几何不规则性与因果有效性之间的关系仍待进一步研究。

原文摘要 · Abstract (English)

Tool-augmented large language models extend their capabilities beyond parametric knowledge through external tools, but tend to invoke them unnecessarily. We investigate whether tool-use decisions have any stable internal representation that can be extracted and manipulated, a question that is non-trivial given that tools exist entirely in context at inference time and have no direct encoding in model weights. We show that steering vectors extracted from heading-anchors positions exert bidirectional causal control over tool-invocation behavior across five open-source models and three domains, suppressing unnecessary tool use most effectively in domains where parametric reasoning suffices. However, geometric analysis reveals that this causal effectiveness does not correspond to clean linear structure: tool-invocation steps exhibit diffuse, bimodal alignment with the suppression vector rather than the consistent negative alignment a linear encoding account would predict, and different tool types recruit largely distinct internal signatures with low cross-tool feature overlap. We hypothesize these geometric properties are indicative of the non-parametric nature of tools, and distinguish tool-use steering vectors from those extracted for parametrically grounded concepts. The relationship between this geometric irregularity and the observed causal effectiveness remains an open question.

工具使用大模型控制机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。