arXiv:2608.13840cs.CLcs.AI2026-08

为生成式AI审计设计可追溯的测量流程,确保结果透明可比。

ASSERT: A Measurement Pipeline for GenAI Audits

论文配图:ASSERT: A Measurement Pipeline for GenAI Audits
图 1 · 摘自论文原文
  • 基于规范定义审计测量标准,确保每项结果可追溯
  • 实测显示不同对话设置使系统评分波动显著
  • 适合需要公正评估AI行为的研究者与监管方

生成式AI系统的审计常以合规率来总结行为表现,研究人员和利益相关者用此率比较系统、追踪退化并决定部署。但报告率不仅反映被测系统,也受测量选择影响,因此率的变化难以判断是系统变化还是测量方式变动所致。我们提出ASSERT——一种基于规范的生成式AI审计测量流程,将每项报告率与生成它的测量选择规范绑定。ASSERT帮助制定行为评分标准与测试用例,执行审计后返回报告率。在对话欺骗行为的案例研究中,我们发现报告率随对话设定、模拟用户、评审者及不合规证据阈值显著变化,这些测量选择会大幅改变报告率,甚至重排系统排名。由于每项率都关联明确规范,不同审计间的差异更易归因与解读。

原文摘要 · Abstract (English)

Audits of generative AI (GenAI) systems often summarize behavior as a reported rate: how often the audited system complies with policy. Researchers and stakeholders use that rate to compare systems, track regressions, and gate deployment. A reported rate reflects both the system under audit and the measurement choices behind it, so a change in the rate can leave it unclear whether the system or those choices moved. We introduce ASSERT, a specification-driven measurement pipeline for GenAI audits that ties each reported rate to a written specification of the measurement choices used to produce it. ASSERT helps draft a behavioral rubric and test cases, then runs the audit against a GenAI system and returns a reported rate. In a case study on conversational deception, we observe that the reported rate moves substantially with the dialogue setup, the simulated user, the judge, and the evidence bar for non-compliance. These measurement choices substantially change the reported rate and can reorder GenAI system rankings. Because each reported rate is tied to an explicit specification, differences across audits are easier to attribute and interpret.

AI审计生成式AI可解释性测评规范

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。