arXiv:2608.16190cs.CRcs.LG2026-08

模型能力比来源多样更能提升可信监控集成效果

Decorrelation Is Not Complementarity: Skill, Not Lineage, Governs Trusted-Monitor Ensembles

  • 用检测技能而非预训练谱系构建监控集成
  • 高技能模型组合能持续超越单个最强模型
  • 跨谱系集成无额外收益,多样性不来自来源

可信监控通过一个低成本的可信模型评估更强但不可信模型的行为,且多样化集成在相同成本下优于单一强监控器。以往方法通过最小化成对相关性实现多样性,但其基础模型相同,未明确多样性来源。本文研究24个开放权重监控器,涵盖9种预训练谱系,检测技能跨度达29倍(在含后门代码数据集上,pAUC@10%FPR从0.028到0.803)。用于构建集成的相关性指标无法预测集成性能,因攻击项上的共识由共享可检测信号和特异性误差两部分构成,二者对集成增益影响相反(斯皮尔曼相关系数分别为-0.25和+0.26),总和仅微弱相关(+0.05);该抵消现象在8次评估中出现7次。模型自身技能主导信号响应(+0.53),而误差基本不变(-0.01),故个体技能高度预测其与群体的一致性(斯皮尔曼0.84,n=24,置换检验p<0.0001)。跨谱系集成在匹配能力下无性能提升(置换检验p=0.13),谱系贡献亦不显著(+0.064,p=0.18)。在自建22监控器池中,添加两个监控器后增益仍为+0.104(p=0.037),表明性能上限取决于所选池子本身。面板整体表现随成员平均技能下降而单调降低(k=2时为-0.66,k=3时为-0.70),且无权重选择优于直接选最优单体。六种攻击者模型下,增益结果全部成立,共识与抵消现象在五种中成立。

原文摘要 · Abstract (English)

Trusted monitoring has a cheap, trusted model score a stronger untrusted model's actions, and a diverse ensemble of them beats a single stronger monitor at matched cost. They are built by minimising average pairwise correlation, and that paper's twelve monitors shared one base model, leaving open what supplies the diversity. We study 24 open-weight monitors spanning nine pretraining lineages and a 29x range of detection skill (pAUC at 10 percent FPR, 0.028 to 0.803) on backdoored code. The metric used to build panels does not predict what a panel is for, and we can say why. Agreement on attack items splits into a shared-detectability signal component and an idiosyncratic error component, which predict ensemble gain with opposite sign (Spearman -0.25 and +0.26), so their sum, the metric actually used, predicts it barely at all (+0.05); the cancellation holds in 7 of 8 evaluations. Skill acts on signal (+0.53) while error stays flat (-0.01), which is why a monitor's own skill predicts its agreement with the pool (Spearman 0.84, n = 24, permutation p below 0.0001). Pretraining lineage is the obvious way to buy decorrelation, and it does not pay. At matched member capability, cross-lineage panels detect no better (permutation p = 0.13), and lineage barely moves the metric either (+0.064, p = 0.18). We report that against ourselves: on our own 22-monitor pool the same test read +0.104 at p = 0.037 until two monitors were added. An earlier pool topping out at pAUC 0.23 had already invalidated another analysis. Such a quantity is a property of the pool assembled. Panel gain over the best member falls monotonically with panel skill (-0.66 at k = 2, -0.70 at k = 3), and no correlation-weighted selection beats picking the single best monitor out of sample. Across six attacker models the gain result holds in all six, the agreement and cancellation results in five of six.

可信监控模型集成检测性能多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。