arXiv:2606.20303cs.CV2026-06被引 1

解决联邦学习在手术AI中因过拟合导致的泛化失败问题

GEN-Guard: Correcting Generalization Failures for Deployable Federated Surgical AI

论文配图:GEN-Guard: Correcting Generalization Failures for Deployable Federated Surgical AI
图 1 · 摘自论文原文
  • 通过隔离客户端数据评估检测泛化失效,避免模型选择偏差
  • 利用差异感知蒸馏学习跨机构鲁棒特征修正,提升未知医院表现
  • 适用于需部署到新医院的手术AI系统,显著改善最差机构性能

外科视频人工智能的联邦学习可在不共享敏感数据的前提下实现协作训练。然而,仅依据参与医院验证数据选择“最优”全局模型的标准评估方式,会导致性能泄漏:所选模型过拟合内部联邦数据,无法泛化至未见机构。我们提出GEN-Guard,一种实用的后处理框架,用于检测并纠正联邦手术AI中的泛化失败。该框架包含基于客户端隔离评估(CBE)的泛化检测,以及基于差异感知蒸馏(DAD)的泛化修正,二者均在标准联邦学习收敛后运行,支持对未见环境的零样本适应。我们首次量化了性能泄漏的严重性,发现标准评估下模型选择失败率超过80%。GEN-Guard在两个多中心临床任务上评估:腹腔镜胆囊切除术手术阶段识别与结肠镜息肉分割。在两数据集上,其均稳定纠正失败现象,使联邦内F1得分最高提升2点,未见机构性能最高提升3点,最差机构性能提升3-9点。性能泄漏是联邦手术AI中系统性且此前被忽视的风险,GEN-Guard提供了有效检测与修正方案,显著增强其跨机构鲁棒性与零样本泛化能力,提升真实世界部署可靠性。

原文摘要 · Abstract (English)

Federated Learning (FL) in surgical video AI enables collaborative model training without sharing sensitive data. However, standard evaluation practices - selecting the "best" global model based only on validation data from participating hospitals - can lead to suboptimal deployment choices. We identify this critical failure mode as performance leakage, where the selected model overfits internal federation data and fails to generalize to unseen institutions. We propose GEN-Guard, a practical post-hoc framework to detect and correct generalization failures in federated surgical AI. It integrates Generalization Detection via Client-Blocked Evaluation (CBE), which validates performance on isolated client distributions to prevent performance leakage, and Generalization Correction through Disagreement-Aware Distillation (DAD), which learns adaptive feature-level corrections for cross-institutional robustness. Both components operate after standard FL convergence while providing robust support for zero-shot adaptation to unseen environments. We first quantify the severity of performance leakage, observing Model Selection Failures (MSFs) exceeding 80% under standard evaluation. GEN-Guard is evaluated on two multi-center clinical challenges: surgical phase recognition in laparoscopic cholecystectomy and polyp segmentation in colonoscopy. Across both datasets, GEN-Guard consistently corrects these failures, improving in-federation F1 scores by up to 2 points, unseen-institution performance by up to 3 points, and worst-case institutional performance by 3-9 points. Performance leakage represents a systematic and previously under-recognized risk in federated surgical AI. GEN-Guard provides a practical solution for detecting and correcting such failures. By improving cross-institutional robustness and zero-shot generalization, it strengthens the reliability of FL for real-world surgical deployment.

联邦学习手术AI泛化性零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。