arXiv:2607.02396cs.AIcs.LG2026-07中稿 · the Mechanistic In…

用快速算法秒级定位大模型拒绝回答的隐藏空间。

Fast Multi-dimensional Refusal Subspaces via RFM-AGOP

  • 基于改进的递归特征机算法,高效识别多维拒绝响应空间。
  • 在Qwen系列模型上实现秒级定位,且比其他方法更准确。
  • 适合需要快速安全调控的大模型研究者与应用开发者。

大型语言模型(LLM)中的行为调控与可解释性日益重要。早期研究认为行为沿单一线性方向编码,但近期发现如拒绝回答有害问题等复杂行为存在于多维子空间中。现有提取方法计算成本高,难以应用于生成长推理链的推理型模型。本文通过将递归特征机(RFM)算法与探测引导初始化结合,可在数秒内完成对推理型(Qwen 3)和非推理型(Qwen 2.5)模型中多维拒绝子空间的识别。尽管RFM计算高效,其在消融任务上的表现也优于其他方法。未来工作将进一步探究不同方法发现的子空间关系。若验证成立,RFM可成为大模型子空间提取的低成本、可扩展补充方案。

原文摘要 · Abstract (English)

Steering and monitoring activations in Large Language Models (LLMs) are increasingly used for both safety and interpretability. Early work assumed behaviours are encoded along single linear directions, but recent findings suggest complex behaviours, such as the refusal to answer harmful queries, live in multi-dimensional subspaces. However, existing methods for extracting these subspaces are computationally expensive, which becomes prohibitive on reasoning models who produce long reasoning traces. By adapting the Recursive Feature Machine (RFM) algorithm -- which can be computed efficiently -- with a probe-informed initialization, we are able to identify the multi-dimensional refusal subspace in seconds, on reasoning (Qwen 3) and non-reasoning (Qwen 2.5) models. While RFM allows for faster subspace identification, it also showed better performances on the ablation task than its alternatives. More work is planned to better understand the relations between subspaces found by different methods. If confirmed, RFM could be a cheap and scalable complement to existing subspace-extraction methods in LLMs.

大模型安全拒绝空间高效算法可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。