通过激活传输实现对生成模型的精准控制,提升安全性和可控性。
Controlling Language and Diffusion Models by Transporting Activations
- 基于最优传输理论,动态调节模型激活值以控制生成内容
- 在语言和图像生成中实现毒性抑制、概念注入与风格精细调控
- 方法轻量高效,不损害模型原有能力,适合实际部署
大型生成模型能力不断增强,其可靠性、安全性及潜在滥用问题日益突出。为应对挑战,近期研究提出通过调控模型激活值来引导或阻止特定概念或行为的生成。本文提出激活传输(AcT)框架,基于最优传输理论,统一并扩展了多种激活调控方法。AcT具有模态无关性,能以极低计算开销实现对模型行为的细粒度控制,同时最小化对模型性能的影响。实验表明,该方法在大语言模型(LLMs)中可有效缓解毒性、引入任意概念并提高真实性;在文本到图像扩散模型(T2Is)中,可实现精细风格控制与概念否定。结果验证了其有效性与通用性。
原文摘要 · Abstract (English)
The increasing capabilities of large generative models and their ever more widespread deployment have raised concerns about their reliability, safety, and potential misuse. To address these issues, recent works have proposed to control model generation by steering model activations in order to effectively induce or prevent the emergence of concepts or behaviors in the generated output. In this paper we introduce Activation Transport (AcT), a general framework to steer activations guided by optimal transport theory that generalizes many previous activation-steering works. AcT is modality-agnostic and provides fine-grained control over the model behavior with negligible computational overhead, while minimally impacting model abilities. We experimentally show the effectiveness and versatility of our approach by addressing key challenges in large language models (LLMs) and text-to-image diffusion models (T2Is). For LLMs, we show that AcT can effectively mitigate toxicity, induce arbitrary concepts, and increase their truthfulness. In T2Is, we show how AcT enables fine-grained style control and concept negation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。