首个评估大模型是否符合企业政策的系统框架,发现模型拒绝不合规请求能力极弱。
COMPASS: A Framework for Evaluating Organization-Specific Policy Alignment in LLMs
- 设计针对企业允许/禁止清单的测试框架,覆盖8个行业场景
- 测试5920个查询,模型对违规请求仅能拒绝13%-40%
- 揭示模型在执行禁止性规则时严重不鲁棒,适合企业AI安全团队使用
随着大语言模型在医疗、金融等高风险企业应用中广泛部署,确保其遵守组织特定政策变得至关重要。现有安全评估仅关注普遍性危害,忽略了组织级政策对齐。我们提出COMPASS(公司/组织政策对齐评估)——首个系统性评估大模型是否遵循组织允许列表和禁止列表政策的框架。我们在八个不同行业场景中应用该框架,生成并验证了5,920个查询,通过精心设计的边缘案例测试模型在常规合规性和对抗鲁棒性上的表现。评估七种前沿模型后发现根本性不对称:模型对合法请求处理准确率超过95%,但对对抗性禁止类违规请求的拒绝率仅为13%-40%。结果表明当前大模型缺乏政策关键部署所需的鲁棒性,确立了COMPASS作为组织级AI安全评估的关键工具。
原文摘要 · Abstract (English)
As large language models are deployed in high-stakes enterprise applications, from healthcare to finance, ensuring adherence to organization-specific policies has become essential. Yet existing safety evaluations focus exclusively on universal harms. We present COMPASS (Company/Organization Policy Alignment Assessment), the first systematic framework for evaluating whether LLMs comply with organizational allowlist and denylist policies. We apply COMPASS to eight diverse industry scenarios, generating and validating 5,920 queries that test both routine compliance and adversarial robustness through strategically designed edge cases. Evaluating seven state-of-the-art models, we uncover a fundamental asymmetry: models reliably handle legitimate requests (>95% accuracy) but catastrophically fail at enforcing prohibitions, refusing only 13-40% of adversarial denylist violations. These results demonstrate that current LLMs lack the robustness required for policy-critical deployments, establishing COMPASS as an essential evaluation framework for organizational AI safety.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。