arXiv:2511.10138cs.IR2025-11KDD被引 32

用统一生成模型替代传统广告推荐多阶段流程,提升效果与效率。

GPR: Towards a Generative Pre-trained One-Model Paradigm for Large-Scale Advertising Recommendation

  • 将广告推荐转为端到端生成任务,统一建模用户意图与广告生成。
  • 在微信公众号广告系统中实现GMV和CTCVR显著提升。
  • 适合工业级推荐系统优化、生成式模型落地场景的从业者参考。

作为连接用户与商业内容的智能基础设施,广告推荐系统在数字经济的信息流动与价值创造中起核心作用。然而,现有多阶段推荐系统存在目标错位与误差传播问题,难以实现全局最优;而统一生成式推荐模型又难以满足实际工业应用需求。为此,我们提出GPR(Generative Pre-trained Recommender),首个将广告推荐重新定义为端到端生成任务的一体化模型框架,取代传统的级联范式。为实现GPR,我们引入三项关键创新:一是设计适配广告场景的统一输入方案与分词方法,将广告与自然内容映射至共享的多层次语义ID空间,增强异构数据间的语义对齐与建模一致性;二是提出异构分层解码器(HHD),通过双解码器架构解耦用户意图建模与广告生成,兼顾训练效率与推理灵活性,同时保持强建模能力;三是提出多阶段联合训练策略,融合多标记预测(MTP)、价值感知微调及层级增强策略优化(HEPO)算法,构建完整生成式推荐链路,统一兴趣建模、价值对齐与策略优化。GPR已在腾讯微信公众号广告系统全面部署,显著提升关键业务指标,包括GMV与CTCVR。

原文摘要 · Abstract (English)

As an intelligent infrastructure connecting users with commercial content, advertising recommendation systems play a central role in information flow and value creation within the digital economy. However, existing multi-stage advertising recommendation systems suffer from objective misalignment and error propagation, making it difficult to achieve global optimality, while unified generative recommendation models still struggle to meet the demands of practical industrial applications. To address these issues, we propose GPR (Generative Pre-trained Recommender), the first one-model framework that redefines advertising recommendation as an end-to-end generative task, replacing the traditional cascading paradigm with a unified generative approach. To realize GPR, we introduce three key innovations spanning unified representation, network architecture, and training strategy. First, we design a unified input schema and tokenization method tailored to advertising scenarios, mapping both ads and organic content into a shared multi-level semantic ID space, thereby enhancing semantic alignment and modeling consistency across heterogeneous data. Second, we develop the Heterogeneous Hierarchical Decoder (HHD), a dual-decoder architecture that decouples user intent modeling from ad generation, achieving a balance between training efficiency and inference flexibility while maintaining strong modeling capacity. Finally, we propose a multi-stage joint training strategy that integrates Multi-Token Prediction (MTP), Value-Aware Fine-Tuning and the Hierarchy Enhanced Policy Optimization (HEPO) algorithm, forming a complete generative recommendation pipeline that unifies interest modeling, value alignment, and policy optimization. GPR has been fully deployed in the Tencent Weixin Channels advertising system, delivering significant improvements in key business metrics including GMV and CTCVR.

广告推荐生成模型端到端工业落地

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。