arXiv:2506.16218cs.CV2025-06ICML被引 3

提升视觉语言模型在联邦学习中的分布外鲁棒性

FOCoOp: Enhancing Out-of-Distribution Robustness in Federated Prompt Learning for Vision-Language Models

  • 用全局、局部和分布外提示捕捉客户端数据异构性
  • 通过双层鲁棒优化增强对分布外变化的适应能力
  • 适合关注隐私保护下模型泛化能力的研究者

面向视觉语言模型的联邦提示学习(FPL)是一种在保护数据隐私的前提下协同优化模型的有效方法。然而,现有方法在性能与鲁棒性之间存在权衡,尤其在分布外(OOD)变化场景下表现不佳,且客户端间固有的分布内(ID)数据异构性进一步加剧了这一挑战。为此,本文提出一种联邦分布外感知上下文优化框架(FOCoOp),利用三类提示——分布内全局提示、本地提示和分布外提示——实现类别级与分布级分离,通过双层分布鲁棒优化动态适应分布外变化。同时,借助半不平衡最优传输机制校准全局提示、看似分布外提示与真实分布外提示的一致性,提升客户端间判别一致性。在多个真实世界数据集上的实验表明,该方法能有效捕捉分布式异构分布,并显著增强对多种分布外变化的鲁棒性。项目代码已开源。

原文摘要 · Abstract (English)

Federated prompt learning (FPL) for vision-language models is a powerful approach to collaboratively adapt models across distributed clients while preserving data privacy. However, existing FPL approaches suffer from a trade-off between performance and robustness, particularly in out-of-distribution (OOD) shifts, limiting their reliability in real-world scenarios. The inherent in-distribution (ID) data heterogeneity among different clients makes it more challenging to maintain this trade-off. To fill this gap, we introduce a Federated OOD-aware Context Optimization (FOCoOp) framework, which captures diverse distributions among clients using ID global prompts, local prompts, and OOD prompts. Specifically, FOCoOp leverages three sets of prompts to create both class-level and distribution-level separations, which adapt to OOD shifts through bi-level distributionally robust optimization. Additionally, FOCoOp improves the discrimination consistency among clients, i.e., calibrating global prompts, seemingly OOD prompts, and OOD prompts by semi-unbalanced optimal transport. The extensive experiments on real-world datasets demonstrate that FOCoOp effectively captures decentralized heterogeneous distributions and enhances robustness of different OOD shifts. The project is available at GitHub.

联邦学习视觉语言模型分布外鲁棒性提示学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。