用最优控制精调视频生成,减少有害内容且不损画质。
Activation Steering of Video Generation Models via Reduced-Order Linear Optimal Control

- 将视频生成建模为动态系统,通过低维潜空间反馈调节激活值
- 在安全与画质基准上显著降低有害生成,保持提示忠实度
- 适合关注生成可控性与安全性的人工智能研究者
基于大规模网络数据训练的文本到视频(T2V)模型可能生成不当内容,亟需干预以减少有害输出而不牺牲视觉质量。激活值调控提供了不同于微调和提示过滤的机制化替代方案,但现有T2V调控方法仍受限于粗粒度、非前瞻性的干预,易导致过度调控和内容退化。为此,我们提出潜空间激活线性二次调节器(LA-LQR),一种用于最小侵入式T2V调控的降阶最优控制框架。LA-LQR将T2V推理建模为动态系统,计算闭环反馈干预信号,使激活值向目标特征设定点逼近,同时惩罚不必要的扰动。为使最优控制适用于高维视频激活,我们基于对比提示对将激活投影至低维、任务相关的子空间,估计该潜空间内的局部线性动态,并求解潜空间LQR问题,获得逐时间步和逐层的调控信号。我们提供了理论边界,关联潜空间设定点跟踪与原始激活空间特征控制;并通过实验证实了降阶潜动态的保真度。在概念调控与视频安全基准测试中,相较于基线方法,LA-LQR有效减少了不安全生成,同时保持提示忠实度与视觉质量。
原文摘要 · Abstract (English)
Text-to-video (T2V) models trained on large-scale web data can generate undesired content, motivating interventions that reduce harmful outputs without sacrificing visual quality. Activation steering offers an attractive mechanistic alternative to finetuning and prompt filtering, but existing T2V steering methods remain limited, typically applying coarse, non-anticipative interventions that can lead to oversteering and content degradation. To close this gap, we propose Latent Activation Linear-Quadratic Regulator (LA-LQR), a reduced-order optimal control framework for minimally invasive T2V steering. LA-LQR formulates T2V inference as a dynamical system and computes closed-loop feedback interventions that steer activations toward desired feature setpoints while penalizing unnecessary perturbations. To make optimal control feasible for high-dimensional video activations, we project activations onto a low-dimensional, task-relevant subspace derived from contrastive prompt pairs, estimate local linear dynamics in this latent space, and solve a latent LQR problem to obtain timestep- and layer-specific steering signals. We provide theoretical bounds relating latent setpoint tracking to raw activation-space feature control, and empirically validate the fidelity of the reduced latent dynamics. On concept steering and video safety benchmarks, LA-LQR reduces unsafe generations relative to baselines, while preserving prompt fidelity and visual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。