联邦学习中平均校准失效,新方法按机构加权保覆盖
When Average Calibration Fails: Site-Conditional Federated Conformal Risk Control
- 按机构加权校准风险曲线,避免平均校准导致的覆盖不足
- 在20所医院数据上,最差机构误报率仅超目标7.8个百分点
- 无需传输原始图像或标签,适合医疗等隐私敏感场景
分布无关的风险控制通过在保留数据上校准预测集阈值实现分割保证。在联邦设置中,标准做法将校准分数合并为单一阈值。我们在真实多中心脑肿瘤数据(FeTS-2022,1,251例患者,20所机构)上量化了关键缺陷:简单合并校准保护了平均医院,但在40%的机构中违反覆盖,最差机构的假阴性率超出目标7.8个百分点。我们追溯此失败源于隐藏设计选择:聚合权重隐式决定了谁的覆盖被保护。样本量加权优化患者级有效性,但可能牺牲机构级可靠性;等机构加权在此基准上显著提升机构级可靠性,效率相当,仅需每机构一个标量。我们提出风险曲线收缩作为合理机制:每机构传输其经验风险曲线(G个标量)和一个超参数n0,平滑插值于局部校准与样本量加权合并校准之间。留一机构分析确定n0=19,实现2.7/20违规,扩展2.0倍。直接拉格朗日预算优化因集中风险于脆弱机构而失败;有限样本修正项至关重要,移除后违规次数翻三倍。无任何患者级图像、掩码或体积级得分离开任一机构。
原文摘要 · Abstract (English)
Conformal risk control (CRC) provides distribution-free segmentation guarantees by calibrating a prediction-set threshold on held-out data. In federated deployments, the standard approach pools calibration scores into a single threshold. We quantify, on real multi-institutional brain tumor data (FeTS-2022, 1,251 subjects, 20 institutions), a critical failure: naive pooled CRC protects the average hospital but violates coverage at 40% of individual institutions, with the worst site exceeding the target false-negative rate by 7.8 percentage points. We trace this failure to a hidden design choice: the aggregation weights implicitly determine whose coverage is protected. Sample-size weighting optimizes patient-level validity but can sacrifice institution-level reliability; equal-site weighting improves institution-level reliability on this benchmark at comparable efficiency, using only a single scalar per site. We propose risk-curve shrinkage as a principled mechanism: each site transmits its empirical risk curve (G scalars) and a single hyperparameter n0 smoothly interpolates between site-specific local calibration and sample-size-weighted pooled calibration. Leave-one-site-out sensitivity analysis identifies n0=19, achieving 2.7/20 violations at 2.0x stretch. Direct Lagrangian budget optimization fails by concentrating risk on vulnerable hospitals; the finite-sample correction term is essential: removing it triples violations. No patient-level images, masks, or per-volume scores leave any site.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。