用实时数据持续监控大模型合规性,让系统自己判断是否违规。
Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring

- 通过运行时观测构建连续合规信号,取代静态审计
- 用多个专业评估模型打分,得分决定模型选用
- 发现评估者分歧是监管不确定性的信号,需人工介入
当前AI合规方法将符合标准视为一次性的审计结果,而非生产系统的持续可测量属性。我们指出,这种合规幻觉在欧盟《人工智能法案》要求持续人工监督和检测行为漂移的背景下显然不适用。为此提出‘从度量中治理’原则,即通过运行时可观测性生成连续合规信号,而非依赖静态评估。基于此,我们推出govllm开源框架,采用治理驱动的路由架构,模型选择由累积合规分数决定,而非仅看延迟或成本。核心是一个由多个专项评估模型组成的评审团(针对欧盟AI法案、GDPR、ANSSI、无障碍性等),其评价分歧被重新定义为需人工干预的监管不确定性信号。我们在5个监管准则下,用4个小型语言模型(1.7B-7B参数)对49组标注的提示/响应对进行全本地部署评估,一致率在51.5%(mistral:7b)至69.1%(phi4-mini)之间,无单一模型在所有准则上占优,支持‘评审团式’设计。进一步发现小模型存在三种结构性失败模式及特定位置偏差,导致在不同提问顺序下一致性下降最高达25个百分点。govllm已开源,以支持可复现的AI治理研究。
原文摘要 · Abstract (English)
Current approaches to AI compliance treat conformity as a binary, audit-time verdict rather than a continuous, measurable property of production systems. We argue that this compliance fiction is structurally ill-suited to the requirements of the EU AI Act, which demands ongoing human oversight and the detection of emergent behavioural drift in deployed systems. We introduce governance from metrics, a principle whereby regulatory compliance is derived as a continuous signal from runtime observability rather than from static assessments. Building on this principle, we present govllm, an open-source framework implementing a governance-driven routing architecture in which model selection is determined by accumulated compliance scores rather than by latency or cost alone. Central to our approach is a panel of regulatory judges - LLM evaluators specialised per criterion (EU AI Act, GDPR, ANSSI, accessibility) - whose inter-judge disagreement we reframe not as noise but as a regulatory uncertainty signal warranting human arbitration. We validate this approach through a ground truth corpus of 49 annotated prompt/response pairs across five regulatory criteria, evaluated by four small language models (SLMs, 1.7B-7B parameters) running fully on-premise. Agreement rates range from 51.5% (mistral:7b) to 69.1% (phi4-mini), with no single model dominating across all criteria - empirically motivating the Profile-as-jury design. We further document three structural failure modes in small regulatory judges and a judge-specific position bias that degrades agreement by up to 25 percentage points across three question-order conditions (original, reversed, permuted). govllm is released as open-source software to support reproducible AI governance research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。