arXiv:2605.06652cs.LGcs.AI2026-05

无标注数据时,如何可靠比较大模型安全性?

When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels

论文配图:When No Benchmark Exists: Validating Comparative LLM Safety Scoring Without Ground-Truth Labels
图 1 · 摘自论文原文
  • 提出基准缺失下的安全评分框架,依赖可控对比与稳定性验证
  • 在挪威数据集上实现0.89~1.00的AUROC,目标差异主导方差(η²≈0.52)
  • 适用于公共采购等场景,需报告完整评估细节而非单一排名

许多部署场景在相关语言、领域或监管制度下尚无标签基准,需在无真实标签时比较候选语言模型的安全性。本文将此设定形式化为基准缺失的比较性安全评分,并明确场景审计可作为部署证据的契约条件。评分仅在固定场景包、评分标准、审计员、评判者、采样配置及重跑预算下有效。由于无标签可用,我们以工具有效性链替代真实标签一致性:对受控的安全与破坏对比的响应性、目标驱动方差主导审计与评判者干扰、多次重跑结果的稳定性。我们在SimpleAudit中实现该链,基于挪威安全场景包验证:安全与破坏目标分离的AUROC值介于0.89至1.00之间,目标身份为最主要方差成分(η² ≈ 0.52),严重性分布十次重跑后趋于稳定。将同一链条应用于Petri,表明两者均适用。差异源于链条上游,即声明-契约执行与部署适配。一挪威公共部门采购案例比较Borealis与Gemma 3显示,更安全模型取决于场景类别和风险度量。因此,必须同时报告评分、匹配差异、临界率、不确定性以及使用的审计员与评判者,而非合并为单一排名。

原文摘要 · Abstract (English)

Many deployments must compare candidate language models for safety before a labeled benchmark exists for the relevant language, sector, or regulatory regime. We formalize this setting as benchmarkless comparative safety scoring and specify the contract under which a scenario-based audit can be interpreted as deployment evidence. Scores are valid only under a fixed scenario pack, rubric, auditor, judge, sampling configuration, and rerun budget. Because no labels are available, we replace ground-truth agreement with an instrumental-validity chain: responsiveness to a controlled safe-versus-abliterated contrast, dominance of target-driven variance over auditor and judge artifacts, and stability across reruns. We instantiate the chain in SimpleAudit, a local-first scoring instrument, and validate it on a Norwegian safety pack. Safe and abliterated targets separate with AUROC values between 0.89 and 1.00, target identity is the dominant variance component ($η^2 \approx 0.52$), and severity profiles stabilize by ten reruns. Applying the same chain to Petri shows that it admits both tools. The substantial differences arise upstream of the chain, in claim-contract enforcement and deployment fit. A Norwegian public-sector procurement case comparing Borealis and Gemma 3 demonstrates the resulting evidence in practice: the safer model depends on scenario category and risk measure. Consequently, scores, matched deltas, critical rates, uncertainty, and the auditor and judge used must be reported together rather than collapsed into a single ranking.

大模型安全无监督评估可信评估公共采购

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。