arXiv:2503.16875cs.IRcs.CL2025-03被引 1

用大模型增强联邦跨域点击率预测,兼顾隐私与精度

Federated Cross-Domain Click-Through Rate Prediction With Large Language Model Augmentation

  • 用大模型扩充用户和物品表征,缓解数据稀疏问题
  • 分离领域特有与共享偏好,提升跨域知识迁移效果
  • 动态调节差分隐私噪声,平衡隐私保护与预测性能

在严格隐私约束下准确预测点击率(CTR)面临严峻挑战,尤其当用户-物品交互数据稀疏且跨域碎片化时。传统跨域CTR方法常假设特征空间同质,依赖集中式数据共享,忽视域间差异及隐私协议带来的权衡。本文提出联邦跨域CTR预测框架FedCCTR-LM,通过同步数据增强、表征解耦与自适应隐私保护来应对上述问题。首先,隐私保护增强网络(PrivAugNet)利用大语言模型丰富用户与物品表征并扩展交互序列,缓解数据稀疏与特征不全。其次,基于对比学习的独立域特定变压器(IDST-CL)模块解耦领域特有与共享用户偏好,结合域内对齐(IDRA)与跨域解耦(CDRD)优化嵌入表示,增强跨域知识迁移。最后,自适应本地差分隐私(AdaLDP)机制动态调整噪声注入,实现隐私保障与预测精度的最佳平衡。在四个真实世界数据集上的实证评估表明,FedCCTR-LM显著优于现有基线,在异构联邦环境中提供鲁棒、隐私保护且可泛化的跨域CTR预测能力。

原文摘要 · Abstract (English)

Accurately predicting click-through rates (CTR) under stringent privacy constraints poses profound challenges, particularly when user-item interactions are sparse and fragmented across domains. Conventional cross-domain CTR (CCTR) methods frequently assume homogeneous feature spaces and rely on centralized data sharing, neglecting complex inter-domain discrepancies and the subtle trade-offs imposed by privacy-preserving protocols. Here, we present Federated Cross-Domain CTR Prediction with Large Language Model Augmentation (FedCCTR-LM), a federated framework engineered to address these limitations by synchronizing data augmentation, representation disentanglement, and adaptive privacy protection. Our approach integrates three core innovations. First, the Privacy-Preserving Augmentation Network (PrivAugNet) employs large language models to enrich user and item representations and expand interaction sequences, mitigating data sparsity and feature incompleteness. Second, the Independent Domain-Specific Transformer with Contrastive Learning (IDST-CL) module disentangles domain-specific and shared user preferences, employing intra-domain representation alignment (IDRA) and crossdomain representation disentanglement (CDRD) to refine the learned embeddings and enhance knowledge transfer across domains. Finally, the Adaptive Local Differential Privacy (AdaLDP) mechanism dynamically calibrates noise injection to achieve an optimal balance between rigorous privacy guarantees and predictive accuracy. Empirical evaluations on four real-world datasets demonstrate that FedCCTR-LM substantially outperforms existing baselines, offering robust, privacy-preserving, and generalizable cross-domain CTR prediction in heterogeneous, federated environments.

联邦学习点击率预测大模型隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。