arXiv:2603.00369cs.CL2026-03ACL

构建首个自然语言请求合规性评估基准,提升AI系统安全可控性。

Policy Compliance of User Requests in Natural Language for AI Systems

  • 设计首个包含多样化合规标注的用户请求数据集
  • 多模型对比验证,揭示现有大模型在合规判断上的显著差异
  • 适用于AI安全、合规审查等工业场景的评估需求

考虑一个组织中用户以自然语言向AI系统提交请求,系统通过执行特定任务来响应。本文关注如何确保这些用户请求符合组织制定的多样政策,以保障AI系统的安全与可靠使用。我们首次提出一个基准数据集,其中包含针对一系列政策具有不同合规程度的标注用户请求,该数据集与科技行业的实际应用密切相关。随后,我们利用该基准评估多种大语言模型在不同解决方案下的合规性评估性能。通过对不同模型和方法在性能指标上的差异分析,凸显了本问题的挑战性。

原文摘要 · Abstract (English)

Consider an organization whose users send requests in natural language to an AI system that fulfills them by carrying out specific tasks. In this paper, we consider the problem of ensuring such user requests comply with a list of diverse policies determined by the organization with the purpose of guaranteeing the safe and reliable use of the AI system. We propose, to the best of our knowledge, the first benchmark consisting of annotated user requests of diverse compliance with respect to a list of policies. Our benchmark is related to industrial applications in the technology sector. We then use our benchmark to evaluate the performance of various LLM models on policy compliance assessment under different solution methods. We analyze the differences on performance metrics across the models and solution methods, showcasing the challenging nature of our problem.

AI合规大模型评估自然语言处理安全控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。