首个外科视觉联邦学习基准,揭示视频时序建模关键作用
Federated Learning for Surgical Vision in Appendicitis Classification: Results of the FedSurg EndoVis 2024 Challenge
- 构建首个外科手术视频联邦学习基准,评估跨中心泛化与个性化适配
- 中心化训练在未见中心仅达26.31% F1,去中心化引入额外性能损失
- 时序建模显著优于帧级方法,参数高效微调避免分类器崩溃
开发可泛化的外科AI需多中心数据,但患者隐私限制了数据共享,使联邦学习(FL)成为自然解决方案。然而,复杂时空手术视频数据的FL应用仍缺乏基准测试。我们提出了FedSurg挑战,首个专注于外科视觉中的联邦学习国际基准,基于多中心腹腔镜阑尾切除数据集(Appendix300初步子集)进行验证。三份提交结果评估了对未见中心的泛化能力及中心特定适应性。集中式与群智学习基线分离出任务难度与去中心化对性能的影响。即使所有数据集中,未见中心的F1得分仅为26.31%,而去中心化训练引入额外可分离的性能下降。时序建模成为主导因素:无论聚合策略如何,视频级时空模型始终优于帧级方法。直接本地微调导致不平衡数据下分类器崩溃;结构化个性化联邦学习结合参数高效微调,是更稳健的中心适配路径。通过严谨统计分析刻画当前FL局限,本工作为外科视频分析中鲁棒、隐私保护的AI系统建立方法论基准。
原文摘要 · Abstract (English)
Developing generalizable surgical AI requires multi-institutional data, yet patient privacy constraints preclude direct data sharing, making Federated Learning (FL) a natural candidate solution. The application of FL to complex, spatiotemporal surgical video data remains largely unbenchmarked. We present the FedSurg Challenge, the first international benchmarking initiative dedicated to FL in surgical vision, evaluated as a proof-of-concept on a multi-center laparoscopic appendectomy dataset (preliminary subset of Appendix300). Three submissions were evaluated on generalization to an unseen center and center-specific adaptation. Centralized and Swarm Learning baselines isolate the contributions of task difficulty and decentralization to observed performance. Even with all data pooled centrally, the task achieved only 26.31\% F1-score on the unseen center, while decentralized training introduced an additional, separable performance penalty. Temporal modeling emerges as the dominant architectural factor: video-level spatiotemporal models consistently outperformed frame-level approaches regardless of aggregation strategy. Naive local fine-tuning leads to classifier collapse on imbalanced local data; structured personalized FL with parameter-efficient fine-tuning represents a more principled path toward center-specific adaptation. By characterizing current FL limitations through rigorous statistical analysis, this work establishes a methodological reference point for robust, privacy-preserving AI systems in surgical video analysis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。