arXiv:2508.09036cs.CYcs.AI2025-08被引 1

测试大模型在隐私与AI治理考试中的表现,发现部分模型已超人类专业水平。

Can We Trust AI to Govern AI? Benchmarking LLM Performance on Privacy and AI Governance Exams

  • 用真实行业认证题库闭卷测试10个主流大模型
  • Gemini 2.5 Pro和GPT-5得分超过人类持证标准
  • 为隐私官和合规人员提供AI工具可信度参考

大型语言模型的快速发展引发了职场对新技术能力的广泛关注。针对隐私专业人士,核心问题是这些AI系统能否可靠支持监管合规、隐私项目管理和AI治理。本研究通过国际隐私专业人士协会(IAPP)的CIPP/US、CIPM、CIPT和AIGP等行业标准认证考试,对包括OpenAI、Anthropic、Google DeepMind、Meta和DeepSeek在内的十款领先开源与闭源大模型进行评估。所有模型在闭卷条件下使用官方样题测试,并与IAPP设定的及格线对比。结果显示,Gemini 2.5 Pro和OpenAI GPT-5等前沿模型持续达到甚至超过人类持证专业人士的标准,展现出在隐私法、技术控制和AI治理方面的显著专业能力。研究揭示了当前大模型在特定领域的强项与局限,为隐私官、合规负责人和技术人员评估高风险数据治理场景中AI工具的可用性提供了实用洞见。本文还建立了基于人类评估的人机基准,帮助专业人士应对人工智能发展与监管风险交织的挑战。

原文摘要 · Abstract (English)

The rapid emergence of large language models (LLMs) has raised urgent questions across the modern workforce about this new technology's strengths, weaknesses, and capabilities. For privacy professionals, the question is whether these AI systems can provide reliable support on regulatory compliance, privacy program management, and AI governance. In this study, we evaluate ten leading open and closed LLMs, including models from OpenAI, Anthropic, Google DeepMind, Meta, and DeepSeek, by benchmarking their performance on industry-standard certification exams: CIPP/US, CIPM, CIPT, and AIGP from the International Association of Privacy Professionals (IAPP). Each model was tested using official sample exams in a closed-book setting and compared to IAPP's passing thresholds. Our findings show that several frontier models such as Gemini 2.5 Pro and OpenAI's GPT-5 consistently achieve scores exceeding the standards for professional human certification - demonstrating substantial expertise in privacy law, technical controls, and AI governance. The results highlight both the strengths and domain-specific gaps of current LLMs and offer practical insights for privacy officers, compliance leads, and technologists assessing the readiness of AI tools for high-stakes data governance roles. This paper provides an overview for professionals navigating the intersection of AI advancement and regulatory risk and establishes a machine benchmark based on human-centric evaluations.

AI治理大模型评测隐私合规机器基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。