arXiv:2506.23759eess.IVcs.CV2025-06被引 5

针对手术视频分割,提出联邦学习新框架,提升跨机构协作效果。

Spatio-Temporal Representation Decoupling and Enhancement for Federated Instrument Segmentation in Surgical Videos

  • 分离空间与时间特征,本地训练个性化背景表示
  • 利用合成数据构建显式表征目标,提升模型泛化能力
  • 适合多中心医疗影像协同建模,尤其手术视频分析

在联邦学习(FL)框架下进行手术器械分割具有重要意义,可在不集中数据的前提下实现多中心协作训练。然而,当前针对手术数据的联邦学习研究仍较少,且现有方法未考虑手术场景的固有特性:一是不同场景下解剖背景差异大但器械表征高度相似;二是可通过手术模拟器低成本生成大规模合成数据。本文提出一种新型个性化联邦学习方案——时空表征解耦增强(FedST),在本地与全局训练中巧妙融入手术领域知识以提升分割性能。具体而言,本地训练采用表征分离与协作(RSC)机制,将查询嵌入层私有化训练以编码各自背景,其余参数全局优化以捕捉器械的一致性特征,包含时序层以建模相似运动模式。进一步设计文本引导的通道选择策略,突出站点特异性特征,促进模型适配各站点。在全局服务器端,提出基于合成数据的显式表征量化(SERQ)方法,定义基于合成数据的显式表征目标,以同步模型融合过程中的收敛性,提升模型泛化能力。实验验证了该方法的有效性,并构建了一个新的个性化联邦学习基准。

原文摘要 · Abstract (English)

Surgical instrument segmentation under Federated Learning (FL) is a promising direction, which enables multiple surgical sites to collaboratively train the model without centralizing datasets. However, there exist very limited FL works in surgical data science, and FL methods for other modalities do not consider inherent characteristics in surgical domain: i) different scenarios show diverse anatomical backgrounds while highly similar instrument representation; ii) there exist surgical simulators which promote large-scale synthetic data generation with minimal efforts. In this paper, we propose a novel Personalized FL scheme, Spatio-Temporal Representation Decoupling and Enhancement (FedST), which wisely leverages surgical domain knowledge during both local-site and global-server training to boost segmentation. Concretely, our model embraces a Representation Separation and Cooperation (RSC) mechanism in local-site training, which decouples the query embedding layer to be trained privately, to encode respective backgrounds. Meanwhile, other parameters are optimized globally to capture the consistent representations of instruments, including the temporal layer to capture similar motion patterns. A textual-guided channel selection is further designed to highlight site-specific features, facilitating model adapta tion to each site. Moreover, in global-server training, we propose Synthesis-based Explicit Representation Quantification (SERQ), which defines an explicit representation target based on synthetic data to synchronize the model convergence during fusion for improving model generalization.

联邦学习手术分割表征解耦合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。