通过分析模型内部层间变化,精准识别文本是否偏离领域意图。
Domain Restriction via Multi SAE Layer Transitions

- 利用稀疏自编码器捕捉模型层间动态,提取领域特异性信号。
- 在Gemma-2 2B和9B模型上验证,能有效区分域外输入。
- 方法轻量且可解释,适合需要安全控制的领域应用。
通用大语言模型在特定领域应用中常因域外(OOD)交互而偏离设计意图。现有检测方法将模型视为不可解释的黑箱,忽略输入的内部处理过程。本文表明,层间转换提供了提取领域特异性特征的可行路径。我们提出几种基于稀疏自编码器(SAE)内部动态学习的轻量级方法,能有效区分域外文本。基于SAE表示的层间转换,有助于理解模型对输入的内在演化过程,揭示其决策机制。我们在Gemma-2 2B和9B模型上进行了全面分析与基准测试,结果表明内部过程在捕捉细粒度输入相关特征方面具有显著效能。
原文摘要 · Abstract (English)
The general-purpose nature of Large Language Models (LLMs) presents a significant challenge for domain-specific applications, often leading to out-of-domain (OOD) interactions that undermine the provider's intent. Existing methods for detecting such scenarios treat the LLM as an uninterpretable black box and overlook the internal processing of inputs. In this work we show that layer transitions provide a promising avenue for extracting domain-specific signature. Specifically, we present several lightweight ways of learning on internal dynamics encoded using a sparse autoencoder (SAE) that exhibit great capability in distinguishing OOD texts. Building on top of SAEs representation transitions enables us to better interpret the LLM internal evolution of input processing and shed light on its decisions. We provide a comprehensive analysis of the method and benchmark it with the gemma-2 2B and 9B models. Our results emphasize the efficacy of the internal process in capturing fine-grained input-related details.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。