用稀疏编码器提前检测并控制大模型生成内容的有害概念。
SCAR: Sparse Conditioned Autoencoders for Concept Detection and Steering in LLMs
- 通过稀疏条件自编码器实现对大模型输出的精准概念控制。
- 在不降低生成质量的前提下,可灵活引导或抑制特定概念(如毒性)。
- 适合关注模型安全、伦理对齐的研究者与应用开发者。
大型语言模型(LLMs)虽能生成类人文本,但其输出可能与用户意图不符甚至产生有害内容。本文提出一种新方法——稀疏条件自编码器(SCAR),通过单一训练模块扩展原有大模型,实现生成前对毒性、安全性和写作风格等概念的检测与调控。该方法支持双向可控性(向/远离特定概念),且在标准评估基准上保持原有生成质量。实验验证了其在多种概念上的有效性,为大模型生成内容的安全与伦理部署提供了可靠框架。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have demonstrated remarkable capabilities in generating human-like text, but their output may not be aligned with the user or even produce harmful content. This paper presents a novel approach to detect and steer concepts such as toxicity before generation. We introduce the Sparse Conditioned Autoencoder (SCAR), a single trained module that extends the otherwise untouched LLM. SCAR ensures full steerability, towards and away from concepts (e.g., toxic content), without compromising the quality of the model's text generation on standard evaluation benchmarks. We demonstrate the effective application of our approach through a variety of concepts, including toxicity, safety, and writing style alignment. As such, this work establishes a robust framework for controlling LLM generations, ensuring their ethical and safe deployment in real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。