arXiv:2505.23024cs.LG2025-05IJCAI被引 3

研究联邦学习中视觉与语言提示学习的差异及优化方法。

An Empirical Study of Federated Prompt Learning for Vision Language Model

  • 对比分析视觉提示与语言提示在联邦学习中的表现差异。
  • 发现提示长度和聚合策略对模型鲁棒性有显著影响。
  • 适合关注隐私保护下多模态模型部署的研究者。

视觉语言模型(VLM)在对齐视觉与语言表征方面表现优异,提示学习已成为适应下游任务的关键技术。然而,将提示学习应用于联邦学习(FL)场景仍缺乏深入研究。本文系统考察了在数据异构(包括标签偏移和领域偏移)条件下,语言提示学习(LPT)与视觉提示学习(VPT)的行为差异。通过大量实验评估客户端规模、聚合策略、提示长度等联邦学习与提示配置的影响,以衡量联邦提示学习(FPL)的鲁棒性。此外,针对标签偏移与领域偏移共存的复杂场景,探索了在计算资源允许时联合使用两种提示的增强策略。研究结果为优化联邦设置下的提示学习提供了实用洞见,助力视觉语言模型在隐私保护环境中的更广泛应用。

原文摘要 · Abstract (English)

The Vision Language Model (VLM) excels in aligning vision and language representations, and prompt learning has emerged as a key technique for adapting such models to downstream tasks. However, the application of prompt learning with VLM in federated learning (FL) scenarios remains underexplored. This paper systematically investigates the behavioral differences between language prompt learning (LPT) and vision prompt learning (VPT) under data heterogeneity challenges, including label skew and domain shift. We conduct extensive experiments to evaluate the impact of various FL and prompt configurations, such as client scale, aggregation strategies, and prompt length, to assess the robustness of Federated Prompt Learning (FPL). Furthermore, we explore strategies for enhancing prompt learning in complex scenarios where label skew and domain shift coexist, including leveraging both prompt types when computational resources allow. Our findings offer practical insights into optimizing prompt learning in federated settings, contributing to the broader deployment of VLMs in privacy-preserving environments.

联邦学习提示学习多模态隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。