arXiv:2504.16612cs.CVcs.LG2025-04被引 4

用联邦学习训练内窥镜视觉模型,保护隐私还能达到中心化效果

Federated EndoViT: Pretraining Vision Transformers via Federated Learning on Endoscopic Image Collections

  • 采用联邦学习框架,结合自适应优化提升跨机构数据训练稳定性
  • 在70万张内窥镜图像上预训练,分割与泛化性能接近中心化模型
  • 适合需要保护患者隐私的多中心外科数据协作场景

数据隐私法规限制了手术领域通用基础模型(FMs)的构建,因无法实现多机构数据聚合。本研究探索联邦学习(FL)作为隐私保护方案,用于协同训练稳健的外科基础模型。提出联邦内窥镜视觉变压器(Federated EndoViT, FL-EndoViT),验证了在去中心化手术环境下掩码自编码器(MAE)预训练策略的有效性。为应对严重数据异构性,架构融合自适应尖锐度感知最小化(FedSAM)。在大规模Endo700k数据集上预训练后,与集中式基线模型在场景分割、动作识别和阶段识别等任务上进行对比评估。结果表明,FedSAM对成功预训练至关重要,克服了传统联邦方法的收敛失败问题。所得的FL-EndoViT性能与集中式模型相当,在数据稀缺、高分辨率分割及新手术事件泛化方面表现显著。同时发现,全端到端微调是获得最佳性能的必要条件。结论:自适应优化的联邦学习可作为构建鲁棒、隐私保护型外科基础模型的可行范式。研究为多中心外科数据科学协作提供了可扩展框架,并强调优化器在处理数据异构性中的关键作用。未来工作应探索基于视频的模型以捕捉时空动态。

原文摘要 · Abstract (English)

Purpose: Data privacy regulations hinder the creation of generalizable foundation models (FMs) for surgery by preventing multi-institutional data aggregation. This study investigates federated learning (FL) as a privacy-preserving solution to collaboratively train robust surgical FMs. Methods: We introduce Federated EndoViT (FL-EndoViT), a federated framework that validates the Masked Autoencoder (MAE) pretraining strategy in a decentralized surgical setting. To ensure convergence under severe data heterogeneity, the architecture integrates adaptive Sharpness-Aware Minimization (FedSAM). Pretrained on the large-scale Endo700k dataset, FL-EndoViT is evaluated against a centralized baseline on different tasks including scene segmentation, action recognition, and phase recognition. Results: FedSAM is critical for successful pretraining, overcoming the convergence failures of standard federated methods. The resulting FL-EndoViT performs comparably to its centralized counterpart, with significant advantages in data-scarce, high-resolution segmentation and generalization to new surgical events. We also establish that full, end-to-end fine-tuning is necessary for optimal performance. Conclusion: This work validates FL with adaptive optimization as a viable paradigm for creating robust, privacy-preserving surgical FMs. Our findings provide a scalable framework for collaborative Surgical Data Science and underscore the optimizer's critical role in handling data heterogeneity. Future work should explore video-based models to incorporate spatiotemporal dynamics.

联邦学习视觉模型外科AI隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。