arXiv:2505.19165cs.AI2025-05被引 16

测试大模型在组织权限下的合规能力,发现顶级模型仍严重失准。

OrgAccess: A Benchmark for Role Based Access Control in Organization Scale LLMs

  • 构建40类权限的合成基准,含易/中/难三难度数据集
  • 最难点上GPT-4.1 F1仅0.27,多权限冲突时性能骤降
  • 揭示大模型在复杂规则遵循与组合推理上的根本缺陷

角色权限控制(RBAC)和层级结构是几乎所有组织信息流动与决策的基础。随着大语言模型(LLMs)在企业场景中作为统一知识库与智能助手的潜力日益显现,一个关键却未被充分探索的挑战浮现:这些模型能否可靠理解并遵守组织层级与权限约束?由于真实企业数据和权限策略具有保密性,评估此能力极为困难。本文提出一个合成但具代表性的基准——OrgAccess,包含40种常见组织角色权限,并构建三类数据:4万条简单(1权限)、1万条中等(3权限组合)、2万条困难(5权限组合),用于测试模型在严格遵守层级规则下的响应能力,尤其在存在重叠或冲突权限的情境下。结果表明,即使最先进的大模型也难以维持合规,且在多权限冲突时性能显著下降。特别地,GPT-4.1在最难点上仅达F1分数0.27。这揭示了大模型在复杂规则遵循与组合推理方面存在根本性短板,为评估其在实际结构化环境中的适用性开辟了新范式。

原文摘要 · Abstract (English)

Role-based access control (RBAC) and hierarchical structures are foundational to how information flows and decisions are made within virtually all organizations. As the potential of Large Language Models (LLMs) to serve as unified knowledge repositories and intelligent assistants in enterprise settings becomes increasingly apparent, a critical, yet under explored, challenge emerges: \textit{can these models reliably understand and operate within the complex, often nuanced, constraints imposed by organizational hierarchies and associated permissions?} Evaluating this crucial capability is inherently difficult due to the proprietary and sensitive nature of real-world corporate data and access control policies. We introduce a synthetic yet representative \textbf{OrgAccess} benchmark consisting of 40 distinct types of permissions commonly relevant across different organizational roles and levels. We further create three types of permissions: 40,000 easy (1 permission), 10,000 medium (3-permissions tuple), and 20,000 hard (5-permissions tuple) to test LLMs' ability to accurately assess these permissions and generate responses that strictly adhere to the specified hierarchical rules, particularly in scenarios involving users with overlapping or conflicting permissions. Our findings reveal that even state-of-the-art LLMs struggle significantly to maintain compliance with role-based structures, even with explicit instructions, with their performance degrades further when navigating interactions involving two or more conflicting permissions. Specifically, even \textbf{GPT-4.1 only achieves an F1-Score of 0.27 on our hardest benchmark}. This demonstrates a critical limitation in LLMs' complex rule following and compositional reasoning capabilities beyond standard factual or STEM-based benchmarks, opening up a new paradigm for evaluating their fitness for practical, structured environments.

权限控制大模型评测企业应用规则遵循

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。