用自动化框架检测大模型是否遵守厂商发布的行为规范。
SpecEval: Evaluating Model Adherence to Behavior Specifications
- 通过解析规范生成针对性提示,让模型自评是否合规。
- 发现多家厂商模型存在高达20%的合规缺口。
- 适合关注AI安全与可信性的研究人员和开发者。
开发基础模型的公司会发布行为准则,承诺其模型应遵循特定规范,但目前尚不清楚模型实际是否遵守。尽管OpenAI、Anthropic、Google等厂商已发布详细的行为规范,涵盖安全约束和品质特征,却缺乏系统性审计。本文提出一种自动化框架,通过解析行为声明、生成针对性提示,并利用模型自身作为评判者,评估模型对规范的遵守程度。核心在于检验提供方规范、模型输出与提供方自用评估模型之间的三重一致性,扩展了以往的生成-验证两重一致性。该框架应用于六家厂商的16个模型,覆盖100多个行为条款,发现系统性不一致,部分厂商的合规率差距高达20%。
原文摘要 · Abstract (English)
Companies that develop foundation models publish behavioral guidelines they pledge their models will follow, but it remains unclear if models actually do so. While providers such as OpenAI, Anthropic, and Google have published detailed specifications describing both desired safety constraints and qualitative traits for their models, there has been no systematic audit of adherence to these guidelines. We introduce an automated framework that audits models against their providers specifications by parsing behavioral statements, generating targeted prompts, and using models to judge adherence. Our central focus is on three way consistency between a provider specification, its model outputs, and its own models as judges; an extension of prior two way generator validator consistency. This establishes a necessary baseline: at minimum, a foundation model should consistently satisfy the developer behavioral specifications when judged by the developer evaluator models. We apply our framework to 16 models from six developers across more than 100 behavioral statements, finding systematic inconsistencies including compliance gaps of up to 20 percent across providers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。