arXiv:2605.18858cs.LGcs.AI2026-05中稿 · ProbML 2026

独立校准的模型聚合后可能集体失准,需用VCG机制纠正。

When Individually Calibrated Models Become Collectively Miscalibrated

论文配图:When Individually Calibrated Models Become Collectively Miscalibrated
图 1 · 摘自论文原文
  • 基于贝叶斯得分最优响应的策略交互导致集体失准
  • 真实数据中误报率恶化达7.25倍,相关性越高越严重
  • VCG机制可对齐激励,适合稀疏或对抗性场景

概率预测系统常将多个模型的概率估计聚合为最终决策。普遍假设是:若各模型单独校准,则聚合结果也应校准。我们发现,在多智能体环境下该假设失效:即使无协同,各模型因策略性响应(博弈论中的Brier最优局部反应)也会导致集体失准,尤其当模型在重叠数据上独立训练时。我们证明,在基于Brier分数的聚合下,信念正相关时每个代理的最优报告会系统性低估正类概率,导致价格悖论(PoA)大于1。在标准设置(n=5,成对相关性=0.5,基率=0.3)中,实测假阴性率恶化达7.25倍。相比之下,基于VCG的聚合通过奖励边际贡献对齐激励,实现主导策略激励相容与近最优性能。三组真实数据集(NSL-KDD、UNSW-NB15、Credit Card Fraud)实验表明,VCG具有强鲁棒性,准确率相当,尤其在数据稀疏和对抗场景表现更优,自适应加权还能缓解分布偏移影响。

原文摘要 · Abstract (English)

Probabilistic prediction systems often aggregate probability estimates from multiple models into a single decision. A common assumption is that if each model is individually calibrated, the aggregate prediction will also be well calibrated. We show that this assumption fails in multi-agent settings: individually calibrated predictors can become collectively miscalibrated when their predictions interact strategically, in the game-theoretic sense of Brier-optimal local response, even without deliberate coordination. This phenomenon arises naturally when agents are independently trained on overlapping data. We prove that under Brier-score-based aggregation with positively correlated beliefs, each agent's individually optimal report systematically underestimates the positive-class probability, yielding a Price of Anarchy greater than one whenever Cov(b_i, b_j) > 0. In a canonical setting (n = 5 agents, pairwise correlation = 0.5, base rate = 0.3), the empirically measured PoA in false-negative rate reaches 7.25x. In contrast, VCG-based aggregation aligns incentives by rewarding marginal contribution, achieving dominant-strategy incentive compatibility and near-optimal performance. Experiments on three real-world datasets (NSL-KDD, UNSW-NB15, Credit Card Fraud) show that VCG provides strong robustness while maintaining comparable accuracy. It performs particularly well in data-sparse and adversarial settings, and adaptive weighting further improves performance under distribution shift.

概率校准博弈论聚合机制异常检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。