通过55个子概念探测,发现大模型有害性集中在低秩子空间。
The Geometry of Harmfulness in LLMs through Subconcept Probing
- 为55种有害子概念设计线性探测器,找到激活空间中的可解释方向。
- 有害性整体构成的子空间秩极低,仅靠主方向就能大幅降低危害。
- 仅调整主方向即可近乎消除有害内容,且模型实用性下降极少。
大语言模型(LLMs)的快速发展加剧了对其有害行为的理解与控制需求。本文提出一种多维度框架,用于探测和调控模型内部的有害内容。针对55种不同有害性子概念(如种族仇恨、就业诈骗、武器相关等),我们学习了对应的线性探测器,得到激活空间中55个可解释的方向。这些方向共同构成一个有害性子空间,我们发现该子空间具有显著的低秩特性。随后,我们测试了完全移除该子空间,以及在主方向上进行调控和移除的效果。结果表明,仅在主方向上进行调控即可近乎消除有害性,同时对模型效用的负面影响极小。研究支持了概念子空间作为可扩展视角来理解大模型行为的观点,并为社区提供了可审计、可加固未来语言模型的实际工具。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have intensified the need to understand and reliably curb their harmful behaviours. We introduce a multidimensional framework for probing and steering harmful content in model internals. For each of 55 distinct harmfulness subconcepts (e.g., racial hate, employment scams, weapons), we learn a linear probe, yielding 55 interpretable directions in activation space. Collectively, these directions span a harmfulness subspace that we show is strikingly low-rank. We then test ablation of the entire subspace from model internals, as well as steering and ablation in the subspace's dominant direction. We find that dominant direction steering allows for near elimination of harmfulness with a low decrease in utility. Our findings advance the emerging view that concept subspaces provide a scalable lens on LLM behaviour and offer practical tools for the community to audit and harden future generations of language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。