用视觉语言模型提升无标签数据下的联邦分割性能。
FRIEREN: Federated Learning with Vision-Language Regularization for Segmentation
- 融合视觉语言模型,通过文本嵌入引导分割解码。
- 在无标签客户端上实现强弱伪标签一致性学习,提升鲁棒性。
- 适用于隐私保护的跨域分割场景,适合研究联邦学习新范式。
联邦学习(FL)为语义分割任务提供了一种保护隐私的跨域适应方案,但面临显著的域偏移挑战,尤其在客户端数据无标签时。现有方法大多不切实际地假设远程客户端有标注数据,或未能充分利用现代视觉基础模型(VFMs)的能力。为此,我们提出一个新任务FFREEDG:模型在服务器端使用带标注源数据预训练后,仅通过客户端的无标签数据进行训练,且不再重新访问源数据。针对该任务,我们提出FRIEREN框架,利用视觉语言模型知识,通过基于CLIP的文本嵌入引导视觉-语言解码器,增强语义消歧,并采用弱-强一致性学习策略,实现伪标签上的鲁棒本地训练。在合成到真实及清晰到恶劣天气等基准测试中,本框架表现优异,性能媲美主流域泛化与适应方法,为未来研究设立了强有力基线。
原文摘要 · Abstract (English)
Federeated Learning (FL) offers a privacy-preserving solution for Semantic Segmentation (SS) tasks to adapt to new domains, but faces significant challenges from these domain shifts, particularly when client data is unlabeled. However, most existing FL methods unrealistically assume access to labeled data on remote clients or fail to leverage the power of modern Vision Foundation Models (VFMs). Here, we propose a novel and challenging task, FFREEDG, in which a model is pretrained on a server's labeled source dataset and subsequently trained across clients using only their unlabeled data, without ever re-accessing the source. To solve FFREEDG, we propose FRIEREN, a framework that leverages the knowledge of a VFM by integrating vision and language modalities. Our approach employs a Vision-Language decoder guided by CLIP-based text embeddings to improve semantic disambiguation and uses a weak-to-strong consistency learning strategy for robust local training on pseudo-labels. Our experiments on synthetic-to-real and clear-to-adverse-weather benchmarks demonstrate that our framework effectively tackles this new task, achieving competitive performance against established domain generalization and adaptation methods and setting a strong baseline for future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。