实时自调优框架提升大模型对抗提示攻击的防御能力。
A Real-Time, Self-Tuning Moderator Framework for Adversarial Prompt Detection
- 动态自适应检测,无需重新训练即可应对新攻击
- 在谷歌Gemini模型上有效抵御现代越狱攻击
- 轻量级设计适合大规模部署,不影响正常对话
随着大模型在社会中广泛应用,确保其对齐性对信息安全至关重要。然而,许多现有防御方法难以快速响应新攻击,会降低模型对正常提示的回应质量,或带来显著的可扩展性障碍。为此,我们提出一种实时自调优(RTST)监督框架,可在保持轻量训练开销的同时防御对抗性攻击。通过在谷歌Gemini模型上评估现代高效越狱攻击的防御效果,结果表明:相比传统微调或分类器模型,该自适应、低侵入性的框架在越狱防御上具有明显优势。
原文摘要 · Abstract (English)
Ensuring LLM alignment is critical to information security as AI models become increasingly widespread and integrated in society. Unfortunately, many defenses against adversarial attacks and jailbreaking on LLMs cannot adapt quickly to new attacks, degrade model responses to benign prompts, or introduce significant barriers to scalable implementation. To mitigate these challenges, we introduce a real-time, self-tuning (RTST) moderator framework to defend against adversarial attacks while maintaining a lightweight training footprint. We empirically evaluate its effectiveness using Google's Gemini models against modern, effective jailbreaks. Our results demonstrate the advantages of an adaptive, minimally intrusive framework for jailbreak defense over traditional fine-tuning or classifier models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。