通过稀疏自编码器揭示大模型在分布外数据下的脆弱边界
At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization

- 用稀疏自编码器分析模型内部表征,定位分布外输入引发的错误概念
- 发现输入含细微拼写错误或越狱提示时,模型激活更多错误概念
- 提供推理时诊断工具,适合关注AI安全与鲁棒性的研究者
预训练的Transformer展现出惊人的泛化能力,有时甚至超越训练数据范围。但在真实场景中,模型常遭遇偏离训练分布的意外或对抗性数据,若缺乏应对机制,可靠性与安全性将下降。本文通过系统实验,提出一种机制化框架,精确刻画Transformer鲁棒性的边界。发现分布外输入(包括微小拼写错误和越狱提示)会促使语言模型在其内部激活更多错误概念。利用这一现象,可量化提示中的分布偏移程度,并据此设计机制化的微调策略以增强大模型鲁棒性。本研究将分布外(OOD)概念从输入数据扩展至模型内部计算过程,提出一种推理时的新型诊断方法,是实现人工智能在科学、商业与政府领域安全部署的关键一步。
原文摘要 · Abstract (English)
Pre-trained transformers have demonstrated remarkable generalization abilities, at times extending beyond the scope of their training data. Yet, real-world deployments often face unexpected or adversarial data that diverges from training data distributions. Without explicit mechanisms for handling such shifts, model reliability and safety degrade, urging more disciplined study of out-of-distribution (OOD) settings for transformers. By systematic experiments, we present a mechanistic framework for delineating the precise contours of transformer model robustness. We find that OOD inputs, including subtle typos and jailbreak prompts, drive language models to operate on an increased number of fallacious concepts in their internals. We leverage this device to quantify and understand the degree of distributional shift in prompts, enabling a mechanistically grounded fine-tuning strategy to robustify LLMs. Expanding the very notion of OOD from input data to a model's private computational processes, a new transformer diagnostic at inference time is a critical step toward making AI systems safe for deployment across science, business, and government.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。