arXiv:2606.20668cs.CRcs.AI2026-06中稿 · ICML

首个独立评估大模型监督系统性能的基准,涵盖检测效果、延迟与成本。

BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems

论文配图:BELLS-O: Evaluating the Operational Trade-offs of LLM Supervision Systems
图 1 · 摘自论文原文
  • 构建跨28个系统的统一评测框架,覆盖输入输出过滤与越狱攻击检测。
  • 专用防护模型在内容审核中表现接近顶级大模型,速度更快、成本更低5-10倍。
  • 越狱检测中前沿大模型更优但代价高,适合对安全性要求极高的场景。

LLM监督系统(如输入/输出过滤器和越狱检测器)是防范部署中滥用的关键防线。现有基准常受厂商偏见影响,忽略成本与延迟,且很少对比专用防护与通用大模型的再利用。我们提出BELLS-O(大模型监督系统操作性评估基准),首个独立的操作性评测体系,评估来自17家厂商的28个系统:包括主流专用防护(如LlamaGuard-4、ShieldGemma-2、Lakera Guard)及被用作监督者的前沿通用模型(如GPT-5.4、Claude Sonnet 4.6、Grok-4.1)。在11类危害内容的输入输出监管与13种越狱攻击技术的检测上,综合评估检测率、误报率、延迟与货币成本。数据集由人工设计提示、专家精选样本与质量控制的合成生成构成,并通过重述消除合成数据中的生成指纹。绘制帕累托前沿揭示使用场景依赖的权衡:在内容审核中,专用系统表现优于通用模型——顶尖系统检测率约95%(对标前沿模型94%),误报率≤2%,同时快5-10倍、成本低约10倍;在越狱检测中,前沿模型虽检测率更高、误报更低,但成本高10-50倍,延迟高5-10倍。我们开放基准、框架、排行榜与数据集,为真实部署提供首个无厂商偏见的选型依据。

原文摘要 · Abstract (English)

LLM supervision systems, namely input/output moderation filters and jailbreak detectors, are the primary safeguard against misuse in deployed AI applications, yet existing benchmarks are often vendor-biased, omit cost and latency, and rarely compare specialized guardrails against repurposed generalist LLMs. We present BELLS-O (Benchmark for the Evaluation of LLM Supervision Systems, Operational), the first independent operational benchmark of LLM supervision systems. BELLS-O evaluates 28 systems from 17 providers: every major specialized guardrail (e.g., LlamaGuard-4, ShieldGemma-2, Lakera Guard) and frontier generalists repurposed as supervisors (e.g., GPT-5.4, Claude Sonnet 4.6, Grok-4.1), jointly on detection rate, false-positive rate, latency, and monetary cost. We cover input/output moderation across 11 harm categories and jailbreak detection across 13 attack techniques, using in-house datasets built from handcrafted prompts, expert-curated samples, and quality-controlled synthetic generation. To suppress latent generator fingerprints in synthetic data, every generated sample is paraphrased. Mapping the Pareto frontier reveals use-case-dependent tradeoffs. On content moderation, specialized supervisors are operationally dominant: top systems match frontier LLMs on detection (~95% vs. 94%) at comparably low false-positive rates (<=2%), while running 5-10x faster and ~10x cheaper. On jailbreak detection, the tradeoff shifts: frontier LLMs achieve higher detection and lower false-positive rates but at 10-50x higher cost and 5-10x higher latency. We release the benchmark, framework, leaderboard, and datasets as the first vendor-neutral basis for selecting safeguards under real deployment constraints.

大模型安全监督系统性能评估成本优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。