揭示大模型误拒现象的根源,提出需针对不同任务定制干预方案。
Over-Refusal and Representation Subspaces: A Mechanistic Analysis of Task-Conditioned Refusal in Aligned LLMs
- 发现误拒源于任务相关的隐空间子结构,而非单一全局方向。
- 有害拒绝由单向量主导,而误拒分布在多个任务簇中且维度更高。
- 早期层已可区分两类拒绝,提示需任务特异性干预策略。
经过对有害请求进行拒绝对齐的语言模型,会表现出过量拒绝:即拒绝看似安全但与有害请求相似的指令。一种自然方法是消除全局拒绝方向,使隐藏状态向量远离或靠近有害拒绝样本,但这仅偶然缓解过量拒绝,同时破坏整体拒绝机制。本文分析了两类拒绝的表征几何结构,发现有害拒绝方向具有任务无关性,可由单一全局向量捕获;而过量拒绝方向则具有任务依赖性,存在于良性任务表示簇内,跨任务变化且占据更高维子空间。线性探测表明,两类拒绝在早期变换器层中即具有可区分的表征。这些发现为为何仅靠全局方向消融无法解决过量拒绝提供了机制解释,并确立了任务特异性几何干预的必要性。
原文摘要 · Abstract (English)
Aligned language models that are trained to refuse harmful requests also exhibit over-refusal: they decline safe instructions that seemingly resemble harmful instructions. A natural approach is to ablate the global refusal direction, steering the hidden-state vectors away or towards the harmful-refusal examples, but this corrects over-refusal only incidentally while disrupting the broader refusal mechanism. In this work, we analyse the representational geometry of both refusal types to understand why this happens. We show that harmful-refusal directions are task-agnostic and can be captured by a single global vector, whereas over-refusal directions are task-dependent: they reside within the benign task-representation clusters, vary across tasks, and span a higher-dimensional subspace. Linear probing suggests that the two refusal types are representationally distinct from the early transformer layers. These findings provide a mechanistic explanation of why global direction ablation alone cannot address over-refusal, and establish that task-specific geometric interventions are necessary.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。