联邦学习下个性化融合视觉语言推荐,保护隐私并提升精准度
Federated Vision-Language-Recommendation with Personalized Fusion
- 服务器生成多模态融合视图,客户端用专家混合模型自适应融合
- 在7个数据集上验证,相比基线提升1.8%-5.3%的推荐准确率
- 适合注重用户隐私与个性化推荐的智能系统研发者
将大型预训练视觉语言模型应用于推荐是新兴方向,我们称之为视觉语言推荐(VLR)。在联邦学习框架下实现面向用户的设备端智能VLR,是提升用户隐私保护并提供个性化体验的关键一步。本文提出FedVLR,一种专为用户特定个性化融合视觉语言表征设计的联邦VLR框架。核心是新颖的双层融合机制:服务器端的多视图融合模块首先生成多样化的预融合多模态视图;随后,每个客户端基于用户交互历史,采用用户专属的专家混合机制自适应整合这些视图。该轻量级个性化融合模块高效支持联邦VLR系统实现。所提FedVLR在七个基准数据集上得到验证,性能显著优于基线方法。
原文摘要 · Abstract (English)
Applying large pre-trained Vision-Language Models to recommendation is a burgeoning field, a direction we term Vision-Language-Recommendation (VLR). Bringing VLR to user-oriented on-device intelligence within a federated learning framework is a crucial step for enhancing user privacy and delivering personalized experiences. This paper introduces FedVLR, a federated VLR framework specially designed for user-specific personalized fusion of vision-language representations. At its core is a novel bi-level fusion mechanism: The server-side multi-view fusion module first generates a diverse set of pre-fused multimodal views. Subsequently, each client employs a user-specific mixture-of-expert mechanism to adaptively integrate these views based on individual user interaction history. This designed lightweight personalized fusion module provides an efficient solution to implement a federated VLR system. The effectiveness of our proposed FedVLR has been validated on seven benchmark datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。