arXiv:2506.23706cs.AIcs.CL2025-06被引 12

用可信执行环境实现可验证的AI安全审计,保护模型与数据隐私

Attestable Audits: Verifiable AI Safety Benchmarks Using Trusted Execution Environments

  • 在可信环境内运行审计,确保结果可验证
  • 支持对Llama-3.1模型的合规性验证,保护双方数据隐私
  • 适合关注AI治理与安全合规的机构使用

基准测试是大规模评估AI模型安全性和合规性的关键手段。然而,现有基准通常无法提供可验证的结果,且在模型知识产权和基准数据集保密性方面存在缺陷。本文提出可验证审计(Attestable Audits),通过在可信执行环境(Trusted Execution Environments)中运行,使用户能够验证其与合规AI模型的交互过程。该方法在模型提供方与审计方互不信任的情况下仍能保护敏感数据,解决了近期人工智能治理框架中提出的验证难题。我们构建了原型系统,在典型的审计基准上对Llama-3.1进行了可行性验证。

原文摘要 · Abstract (English)

Benchmarks are important measures to evaluate safety and compliance of AI models at scale. However, they typically do not offer verifiable results and lack confidentiality for model IP and benchmark datasets. We propose Attestable Audits, which run inside Trusted Execution Environments and enable users to verify interaction with a compliant AI model. Our work protects sensitive data even when model provider and auditor do not trust each other. This addresses verification challenges raised in recent AI governance frameworks. We build a prototype demonstrating feasibility on typical audit benchmarks against Llama-3.1.

AI安全可信计算审计验证

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。