不需重新训练,用注意力重加权就能让AI避开生成有害内容。
Attention Shift: Steering AI Away from Unsafe Content
- 通过调整注意力权重,动态抑制模型中的有害概念。
- 在直接和对抗性越狱提示下,均有效降低有害内容生成率。
- 适合需要快速部署安全机制的AI系统开发者。
本研究探讨了先进生成模型中产生有害内容的问题,提出一种无需额外训练的新型方法:通过注意力重加权,在推理阶段移除有害概念。我们对比了该方法与现有消融方法的性能,采用定性和定量指标评估其在直接及对抗性越狱提示下的表现。分析了实验结果可能的原因,并讨论了内容限制技术的局限性与更广泛影响。
原文摘要 · Abstract (English)
This study investigates the generation of unsafe or harmful content in state-of-the-art generative models, focusing on methods for restricting such generations. We introduce a novel training-free approach using attention reweighing to remove unsafe concepts without additional training during inference. We compare our method against existing ablation methods, evaluating the performance on both, direct and adversarial jailbreak prompts, using qualitative and quantitative metrics. We hypothesize potential reasons for the observed results and discuss the limitations and broader implications of content restriction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。