arXiv:2603.04837cs.AI2026-03

为大模型设计可审计的实时行为治理框架,显著降低风险暴露率。

Design Behaviour Codes (DBCs): A Taxonomy-Driven Layered Governance Benchmark for Large Language Models

  • 构建分层治理框架,在推理阶段通过系统提示控制模型行为。
  • 风险暴露率从7.19%降至4.55%,相对减少36.8%,优于传统安全提示。
  • 支持多领域合规评估,适合监管机构与模型开发者使用。

我们提出动态行为约束(DBC)基准,首个用于评估结构化150项控制行为治理层(MDBC系统)在推理阶段对大语言模型有效性实证框架。不同于训练时对齐方法或事后内容审核接口,DBC是模型无关、可映射司法辖区且可审计的系统级提示治理层。我们在包含30个领域的风险分类体系中,按六个集群(幻觉与校准、偏见与公平性、恶意使用、隐私与数据保护、鲁棒性与可靠性、错位代理)进行评估,采用五种对抗攻击策略(直接、角色扮演、少样本、假设性、权威伪造)测试三类模型家族。三臂对照设计(基础模型、基础+安全提示、基础+DBC)实现风险降低的因果归因。关键结果:DBC将总体风险暴露率(RER)从7.19%(基础)降至4.55%,相对减少36.8%,而标准安全提示仅降低0.6%。MDBC遵从度评分从8.6/10提升至8.7/10。欧盟人工智能法案自动化评分达8.5/10。三评审员评估组显示Fleiss kappa > 0.70(强一致性),验证自动流程可靠性。聚类消融分析表明‘完整性保护’子集(MDBC 081 099)贡献最大风险下降,灰盒攻击绕过率为4.83%。我们开源全部代码、提示库与评估工具,支持复现与模型演进追踪。

原文摘要 · Abstract (English)

We introduce the Dynamic Behavioral Constraint (DBC) benchmark, the first empirical framework for evaluating the efficacy of a structured, 150-control behavioral governance layer, the MDBC (Madan DBC) system, applied at inference time to large language models (LLMs). Unlike training time alignment methods (RLHF, DPO) or post-hoc content moderation APIs, DBCs constitute a system prompt level governance layer that is model-agnostic, jurisdiction-mappable, and auditable. We evaluate the DBC Framework across a 30 domain risk taxonomy organized into six clusters (Hallucination and Calibration, Bias and Fairness, Malicious Use, Privacy and Data Protection, Robustness and Reliability, and Misalignment Agency) using an agentic red-team protocol with five adversarial attack strategies (Direct, Roleplay, Few-Shot, Hypothetical, Authority Spoof) across 3 model families. Our three-arm controlled design (Base, Base plus Moderation, Base plus DBC) enables causal attribution of risk reduction. Key findings: the DBC layer reduces the aggregate Risk Exposure Rate (RER) from 7.19 percent (Base) to 4.55 percent (Base plus DBC), representing a 36.8 percent relative risk reduction, compared with 0.6 percent for a standard safety moderation prompt. MDBC Adherence Scores improve from 8.6 by 10 (Base) to 8.7 by 10 (Base plus DBC). EU AI Act compliance (automated scoring) reaches 8.5by 10 under the DBC layer. A three judge evaluation ensemble yields Fleiss kappa greater than 0.70 (substantial agreement), validating our automated pipeline. Cluster ablation identifies the Integrity Protection cluster (MDBC 081 099) as delivering the highest per domain risk reduction, while graybox adversarial attacks achieve a DBC Bypass Rate of 4.83 percent . We release the benchmark code, prompt database, and all evaluation artefacts to enable reproducibility and longitudinal tracking as models evolve.

大模型治理行为控制风险评估合规性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。