动态调节差分隐私噪声,让联邦学习模型既安全又可解释。
Pushing the (Decision) Boundaries: Dynamically Calibrating Differentially Private Noise to Explainability in Federated Learning

- 根据预测置信度、决策边界敏感度和注意力集中度,实时调整隐私噪声。
- 在医学影像数据上,解释力提升最高达5倍,准确率提高超10%。
- 适合医疗等需要可信解释的高风险场景,兼顾隐私与可解释性。
联邦学习结合差分隐私(DP)正被广泛用于保护分布式机器学习中的数据隐私。然而,DP噪声会扭曲模型表示,降低解释质量,限制了在临床辅助诊断等需可信解释场景中的应用。现有方法仅用静态特征重要性信号调节噪声,只能事后分析,无法在训练中动态优化解释效果。本文提出XCal-FL,一种闭环的、以可解释性驱动的本地训练算法,适用于跨站点图像分类任务。该方法基于三个互补信号动态校准DP噪声:(1) 预测置信度变化,衡量对模型信心的因果影响;(2) 反事实边界间距,捕捉决策边界的敏感性;(3) 注意力集中度,量化模型关注区域的空间一致性,并通过自适应隐私会计保障形式化差分隐私。在三个医学影像数据集上,不同联邦配置下的实验表明,XCal-FL生成的全局模型更准确、更可解释,预测性能提升超过10%,解释力最高提升5倍,优于当前最先进的自适应差分隐私方法。同时,该方法具有更高的隐私预算效率,单位累积隐私损失带来的准确率与解释力增益更大。分析还发现,解释力与隐私损失呈非线性关系,而预测性能近似线性增长,说明可解释性是独立于泛化性能的隐私权衡维度,对关键决策场景的训练与隐私预算分配具有重要启示。
原文摘要 · Abstract (English)
Federated Learning (FL) with Differential Privacy (DP) is increasingly adopted to preserve data confidentiality in distributed machine learning. However, DP noise distorts learned representations and degrades explanation fidelity, limiting differentially private FL where trustworthy explanations are required, such as assistive clinical diagnosis. Prior work adapted DP noise with static feature-importance signals, restricting explainability to post hoc analysis and precluding noise calibration to explanation quality during training. We propose XCal-FL, a closed-loop, explainability-driven local training algorithm for image classification in cross-silo FL that dynamically calibrates DP noise from three complementary signals: (1) prediction logit variations, measuring causal influence on model confidence, (2) counterfactual margins, capturing decision-boundary sensitivity, and (3) saliency concentration, quantifying spatial coherence of model attention, while enforcing formal DP guarantees via adaptive privacy accounting. Experiments on three medical imaging datasets across varying FL configurations show that XCal-FL yields more accurate and interpretable global models, improving predictive performance by over 10\% and explanation fidelity by up to 5$\times$ over static-noise FL, and outperforming state-of-the-art adaptive DP methods in fidelity. XCal-FL also achieves higher privacy-budget efficiency, turning each unit of cumulative privacy loss into larger gains in both accuracy and explanation fidelity. Our analysis further reveals that, unlike predictive performance, which scales roughly linearly with privacy loss, explanation fidelity exhibits non-linear dynamics. These findings suggest explainability is a distinct dimension of the privacy trade-off that cannot be inferred from utility alone, with implications for training and privacy-budget allocation in decision-critical applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。