arXiv:2604.09213cs.CV2026-04

通过操控中间激活值,实现对扩散模型生成内容的精准控制。

SHIFT: Steering Hidden Intermediates in Flow Transformers

论文配图:SHIFT: Steering Hidden Intermediates in Flow Transformers
图 1 · 摘自论文原文
  • 在推理时动态调整特定层和时间步的中间激活值
  • 可有效移除或改变目标视觉概念,保持图像质量
  • 无需重训练,适用于多种提示和风格迁移

扩散模型已成为高质量图像生成的主流方法。基于DiT的扩散模型尤其在遵循提示方面表现优异。我们提出SHIFT,一种简单而高效的轻量级框架,通过在推理阶段对中间激活值进行定向操控,实现对DiT扩散模型中特定概念的移除。该方法受大型语言模型中激活值操控的启发,学习可动态应用的引导向量,作用于选定层与时间步,以抑制不需要的视觉概念,同时保留提示中的其余内容和整体图像质量。除了概念抑制,同一机制还可将生成结果转向指定风格域,或引导添加/更改目标对象。实验表明,SHIFT能在不进行耗时重训练的情况下,对DiT生成过程提供高效且灵活的控制,适用于多样化的提示与目标。

原文摘要 · Abstract (English)

Diffusion models have become leading approaches for high-fidelity image generation. Recent DiT-based diffusion models, in particular, achieve strong prompt adherence while producing high-quality samples. We propose SHIFT, a simple but effective and lightweight framework for concept removal in DiT diffusion models via targeted manipulation of intermediate activations at inference time, inspired by activation steering in large language models. SHIFT learns steering vectors that are dynamically applied to selected layers and timesteps to suppress unwanted visual concepts while preserving the prompt's remaining content and overall image quality. Beyond suppression, the same mechanism can shift generations into a desired \emph{style domain} or bias samples toward adding or changing target objects. We demonstrate that SHIFT provides effective and flexible control over DiT generation across diverse prompts and targets without time-consuming retraining.

扩散模型图像生成可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。