联邦学习中生成能捕捉风格差异的视觉提示,提升跨类别泛化能力。
Federated Cross-Modal Style-Aware Prompt Generation
- 融合多尺度视觉特征与客户端风格统计生成提示
- 在多个数据集上准确率超越现有联邦提示方法
- 适合处理非独立同分布和风格多样的分布式数据
提示学习已推动CLIP等视觉-语言模型在多种任务中表现优异,因其计算效率高,非常适合联邦学习。然而,传统仅依赖最终层特征的方法忽略了去中心化客户端数据中的丰富多尺度视觉线索和领域特定风格变化。为此,我们提出FedCSAP(联邦跨模态风格感知提示生成)。该框架利用CLIP视觉编码器的低、中、高层特征,并结合从批量统计中提取的客户端特定风格指标。通过融合复杂视觉细节与文本上下文,FedCSAP生成鲁棒、上下文感知的提示词元,兼具区分性与非冗余性,显著提升对已见与未见类别的泛化能力。在联邦学习范式下,该方法通过本地训练与全局聚合保障数据隐私,有效应对非独立同分布类别分布及多样化领域风格。在多个图像分类数据集上的全面实验表明,FedCSAP在准确率与整体泛化性能上均优于现有联邦提示学习方法。
原文摘要 · Abstract (English)
Prompt learning has propelled vision-language models like CLIP to excel in diverse tasks, making them ideal for federated learning due to computational efficiency. However, conventional approaches that rely solely on final-layer features miss out on rich multi-scale visual cues and domain-specific style variations in decentralized client data. To bridge this gap, we introduce FedCSAP (Federated Cross-Modal Style-Aware Prompt Generation). Our framework harnesses low, mid, and high-level features from CLIP's vision encoder alongside client-specific style indicators derived from batch-level statistics. By merging intricate visual details with textual context, FedCSAP produces robust, context-aware prompt tokens that are both distinct and non-redundant, thereby boosting generalization across seen and unseen classes. Operating within a federated learning paradigm, our approach ensures data privacy through local training and global aggregation, adeptly handling non-IID class distributions and diverse domain-specific styles. Comprehensive experiments on multiple image classification datasets confirm that FedCSAP outperforms existing federated prompt learning methods in both accuracy and overall generalization.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。