arXiv:2504.00441cs.CRcs.AI2025-04被引 14

研究发现安全防护越强,模型越难用,无法兼顾两者。

No Free Lunch with Guardrails

  • 构建评估框架,量化安全、风险与可用性之间的权衡
  • 测试多款主流防护系统,发现强度提升必然降低实用性
  • 提出设计优化蓝图,帮助平衡安全性与用户体验

随着大语言模型和生成式AI广泛应用,防护机制成为保障安全使用的关键工具。然而,强化防护往往带来可用性下降,而过度灵活又可能暴露于攻击风险。本文通过构建评估框架,系统衡量不同防护机制在风险、安全与可用性间的平衡关系,并开发一种高效防护方案。实验评估了Azure Content Safety、Bedrock Guardrails、OpenAI Moderation API、Guardrails AI、Nemo Guardrails及Enkrypt AI等主流防护系统,同时测试GPT-4o、Gemini 2.0-Flash、Claude 3.5-Sonnet和Mistral Large-Latest在简单提示、详细提示及带思维链(CoT)提示下的表现。结果表明:安全与可用性之间存在根本权衡,不存在免费的午餐。研究提出了优化防护设计的蓝图,以在最小化风险的同时保持实用性能。

原文摘要 · Abstract (English)

As large language models (LLMs) and generative AI become widely adopted, guardrails have emerged as a key tool to ensure their safe use. However, adding guardrails isn't without tradeoffs; stronger security measures can reduce usability, while more flexible systems may leave gaps for adversarial attacks. In this work, we explore whether current guardrails effectively prevent misuse while maintaining practical utility. We introduce a framework to evaluate these tradeoffs, measuring how different guardrails balance risk, security, and usability, and build an efficient guardrail. Our findings confirm that there is no free lunch with guardrails; strengthening security often comes at the cost of usability. To address this, we propose a blueprint for designing better guardrails that minimize risk while maintaining usability. We evaluate various industry guardrails, including Azure Content Safety, Bedrock Guardrails, OpenAI's Moderation API, Guardrails AI, Nemo Guardrails, and Enkrypt AI guardrails. Additionally, we assess how LLMs like GPT-4o, Gemini 2.0-Flash, Claude 3.5-Sonnet, and Mistral Large-Latest respond under different system prompts, including simple prompts, detailed prompts, and detailed prompts with chain-of-thought (CoT) reasoning. Our study provides a clear comparison of how different guardrails perform, highlighting the challenges in balancing security and usability.

安全防护大模型可用性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。