arXiv:2505.15805cs.CL2025-05EMNLP被引 6

测试大模型在问答中遵守安全政策的能力,发现其易受间接攻击泄露敏感信息。

Keep Security! Benchmarking Security Policy Preservation in Large Language Model Contexts Against Indirect Attacks in Question Answering

  • 构建真实场景下的安全政策测试集CoPriva,包含直接与间接攻击
  • 10个大模型中多数违反用户政策,间接攻击下泄露率显著上升
  • 模型能答对问题但难遵守规则,需显式提示才可修正输出

随着大语言模型(LLMs)在企业与政府等敏感领域日益广泛应用,确保其在上下文中遵循用户定义的安全策略(尤其是信息不披露)至关重要。现有研究多关注通用安全与社会敏感数据,但缺乏针对上下文安全保护的大规模基准测试。为此,我们提出新型大规模基准数据集CoPriva,用于评估LLM在问答任务中对上下文非披露政策的遵守情况。该数据集基于真实场景构建,包含明确的安全政策及旨在获取禁止信息的直接与复杂间接攻击性问题。我们在10个LLM上进行评估,发现模型普遍存在严重漏洞:多数模型会违背用户设定的安全政策,导致敏感信息泄露。这一问题在面对间接攻击时尤为突出,暴露出当前大模型安全对齐在敏感应用中的关键缺陷。分析表明,尽管模型通常能正确回答问题,但在生成过程中难以融入策略约束;而当被显式提示时,表现出部分修正能力。研究强调亟需更稳健的方法以保障上下文安全性。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) are increasingly deployed in sensitive domains such as enterprise and government, ensuring that they adhere to user-defined security policies within context is critical-especially with respect to information non-disclosure. While prior LLM studies have focused on general safety and socially sensitive data, large-scale benchmarks for contextual security preservation against attacks remain lacking. To address this, we introduce a novel large-scale benchmark dataset, CoPriva, evaluating LLM adherence to contextual non-disclosure policies in question answering. Derived from realistic contexts, our dataset includes explicit policies and queries designed as direct and challenging indirect attacks seeking prohibited information. We evaluate 10 LLMs on our benchmark and reveal a significant vulnerability: many models violate user-defined policies and leak sensitive information. This failure is particularly severe against indirect attacks, highlighting a critical gap in current LLM safety alignment for sensitive applications. Our analysis reveals that while models can often identify the correct answer to a query, they struggle to incorporate policy constraints during generation. In contrast, they exhibit a partial ability to revise outputs when explicitly prompted. Our findings underscore the urgent need for more robust methods to guarantee contextual security.

安全对齐政策遵守大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。