arXiv:2502.12197cs.CLcs.AI2025-02被引 40

提升大模型对系统指令的鲁棒性,防止用户对抗时失效

A Closer Look at System Prompt Robustness

  • 基于真实场景数据构建评估与微调集,模拟复杂交互
  • 引入无分类器引导等推理干预,显著提升指令遵循能力
  • 适合关注大模型安全可控部署的研究者与开发者

系统提示已成为控制大语言模型在对话和智能体场景中行为的关键接口。开发者依赖系统提示定义重要上下文、输出格式、个性特征、安全约束、内容策略及防护措施,这些都要求模型在面对冲突或对抗性用户输入时仍能稳健遵循系统指令。然而实践中,模型常忽略相关安全规则或无法协调系统与用户的矛盾需求。本文通过从OpenAI GPT Store和HuggingFace HuggingChat收集的真实提示,构建了新的评估与微调数据集,研究多种提升系统提示鲁棒性的方法。实验表明,使用真实微调数据结合推理阶段的无分类器引导等干预手段,可显著改善模型表现。我们还分析了OpenAI与DeepSeek最新推理模型在该任务上的表现,发现虽有进步但效果不均。总体而言,现有技术仍不足以保障系统提示的鲁棒性,亟需进一步研究。

原文摘要 · Abstract (English)

System prompts have emerged as a critical control surface for specifying the behavior of LLMs in chat and agent settings. Developers depend on system prompts to specify important context, output format, personalities, guardrails, content policies, and safety countermeasures, all of which require models to robustly adhere to the system prompt, especially when facing conflicting or adversarial user inputs. In practice, models often forget to consider relevant guardrails or fail to resolve conflicting demands between the system and the user. In this work, we study various methods for improving system prompt robustness by creating realistic new evaluation and fine-tuning datasets based on prompts collected from from OpenAI's GPT Store and HuggingFace's HuggingChat. Our experiments assessing models with a panel of new and existing benchmarks show that performance can be considerably improved with realistic fine-tuning data, as well as inference-time interventions such as classifier-free guidance. Finally, we analyze the results of recently released reasoning models from OpenAI and DeepSeek, which show exciting but uneven improvements on the benchmarks we study. Overall, current techniques fall short of ensuring system prompt robustness and further study is warranted.

大模型安全提示鲁棒性指令遵循

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。