让提示词同时变聪明又更安全,用多目标进化算法优化。
Survival of the Safest: Towards Secure Prompt Optimization through Interleaved Multi-Objective Evolution
- 用语义、反馈和交叉变异混合策略演化提示词
- 在多个数据集上表现优异且显著提升安全性
- 适合需要兼顾性能与安全的工业级应用
大型语言模型(LLMs)展现出强大能力,但其提示词优化长期侧重性能而忽视安全。为此,我们提出「最安全者生存」(SoS)——一种新型多目标提示词优化框架,可同步提升性能与安全性。SoS采用交错式多目标进化策略,融合语义、反馈与交叉变异机制,高效探索高维离散提示空间。相比计算开销大的帕累托前沿方法,SoS具备良好可扩展性,保持低计算成本。该方法支持灵活的目标权重设置,生成多样化优化候选提示,使用户可根据实际需求选择最优方案。在多个基准数据集上的实验验证表明,相比单目标方法,SoS在保证高性能的同时显著增强安全性和鲁棒性。这一进展推动了高绩效且安全的LLM系统在工业场景中的部署。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable capabilities; however, the optimization of their prompts has historically prioritized performance metrics at the expense of crucial safety and security considerations. To overcome this shortcoming, we introduce "Survival of the Safest" (SoS), an innovative multi-objective prompt optimization framework that enhances both performance and security in LLMs simultaneously. SoS utilizes an interleaved multi-objective evolution strategy, integrating semantic, feedback, and crossover mutations to effectively traverse the prompt landscape. Differing from the computationally demanding Pareto front methods, SoS provides a scalable solution that expedites optimization in complex, high-dimensional discrete search spaces while keeping computational demands low. Our approach accommodates flexible weighting of objectives and generates a pool of optimized candidates, empowering users to select prompts that optimally meet their specific performance and security needs. Experimental evaluations across diverse benchmark datasets affirm SoS's efficacy in delivering high performance and notably enhancing safety and security compared to single-objective methods. This advancement marks a significant stride towards the deployment of LLM systems that are both high-performing and secure across varied industrial applications
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。