为基于原则的监管裁判大模型设计四维可信度评测基准
A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation
- 构建涵盖准确率、改写鲁棒性、对抗鲁棒性和校准性的四轴评测体系
- 120B大模型在恶意注入场景下准确率从0.74暴跌至0.27,暴露合规表演问题
- 提出可审计的校准评估方法Ceca,支持逐例反事实归因分析
基于原则的监管(如“公平、清晰且不误导”或“实现良好结果”)无法简化为二元判断,大模型作为裁判的应用日益广泛。我们主张此类裁判必须在四个维度上进行评估:准确性、改写鲁棒性、对抗鲁棒性和校准性。为此,我们发布了Principle-Bench,包含168个加密资产金融宣传场景,映射到英国金融行为监管局(FCA)的两项原则,通过预注册标准生成改写、对抗关键词填充和边界扰动样本;这是首个覆盖全部四轴的基于原则监管评测基准。我们还引入了Ceca(校准的范例聚类评估):一种可校准、可审计的评估器,可输出每个示例的精确反事实归因。在关键词计数、三种句子嵌入模型、一个开源权重大模型裁判及校准级联中,无任何方法在所有四轴上占优。120B大模型在良性输入上表现最强,但在关键词填充的消费者保护条款输入上准确率下降47个百分点(从0.74降至0.27),体现“合规表演”现象。来自不同模型家族的另一裁判在此子集上仅达成0.16的科恩κ值,将失败定位至模型本身而非语料库。任何部署级的大模型裁判用于基于原则的监管,都必须报告各原则下的对抗欺骗情况与事后校准结果。
原文摘要 · Abstract (English)
Principle-based regulation, with evaluative standards such as "fair, clear, and not misleading" or "deliver good outcomes", cannot be reduced to binary predicates, and LLM-as-judge is increasingly used as the substitute. Our position is that any such judge must be evaluated on four axes: accuracy, paraphrase robustness, adversarial robustness, and calibration. We release Principle-Bench, 168 cryptoasset financial-promotion scenarios mapped to two UK FCA principles, with paraphrase, adversarial keyword-stuffing, and boundary perturbations authored under a pre-registered rubric; the first benchmark covering all four axes for principle-based regulation. We also introduce Ceca (Calibrated Exemplar-Cluster Assessment): a calibrated, auditable assessor that emits exact per-exemplar counterfactual attributions. Across keyword counting, three sentence-transformer embedders, an open-weight LLM-judge, and a calibrated cascade, no method dominates all four axes. A 120B LLM-judge, strongest on benign inputs, loses 47 accuracy points (0.74 to 0.27) on keyword-stuffed Consumer Duty inputs: "compliance theatre." A second judge from a different model family agrees only at Cohen's kappa = 0.16 on that split, localising the failure to the model rather than the corpus. Any deployment-grade LLM-judge for principle-based regulation must report per-principle adversarial deception and post-hoc calibration alongside aggregate accuracy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。