arXiv:2605.05220cs.LGcs.AI2026-05

提出一种最小扰动的概念操控框架,提升生成模型控制的理论严谨性与灵活性。

MidSteer: Optimal Affine Framework for Steering Generative Models

论文配图:MidSteer: Optimal Affine Framework for Steering Generative Models
图 1 · 摘自论文原文
  • 基于仿射变换构建概念操控的统一理论,将消解不良行为视为特例。
  • 提出LEACE-Switch方法,在特定条件下实现最优仿射切换。
  • 新框架MidSteer放宽假设,适用于视觉扩散模型与大语言模型等多场景。

通过中间表示的引导,已成为控制生成模型的有效策略,尤其在部署后对齐与安全场景中表现突出。然而,尽管该方法在实践中取得成功,尚缺乏系统的理论框架。本文首次建立引导与仿射概念擦除之间的联系,证明标准去除不良行为的方法是LEACE(一种闭式仿射擦除方法)的特例。随后,我们提出了一个原理性的概念切换框架LEACE-Switch,刻画其在何种假设下可提供最优仿射解。在此基础上,进一步提出更通用的仿射框架MidSteer(最小扰动概念引导),放宽原有假设,支持定向且最小扰动的转换。实验表明,MidSteer在多种任务、模态和架构(包括视觉扩散模型与大语言模型)中均表现优异。

原文摘要 · Abstract (English)

Steering intermediate representations has emerged as a powerful strategy for controlling generative models, particularly in post-deployment alignment and safety settings. However, despite its empirical success, it currently lacks a comprehensive theoretical framework. In this paper, we bridge this gap by formalizing the theory of concept steering. First, we establish a link between steering and affine concept erasure, proving that the standard approach for removing unwanted behaviors is a special case of LEACE (a closed-form method for affine erasure). Next, we formulate a principled theoretical framework for concept switching, LEACE-Switch, and characterize the assumptions under which it provides an optimal affine solution. Building on this analysis, we then introduce MidSteer (Minimal Disturbance concept Steering), a more general affine framework for concept manipulation that relaxes these assumptions and enables directed, minimal-disturbance transformations. We demonstrate that MidSteer performs favorably across a range of tasks, modalities, and architectures, including vision diffusion models and large language models.

生成模型概念引导仿射变换

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。