GALA通过生成式强化学习对齐多模态表征,提升淘宝闪购推荐效果。
GALA: Generative Aligned Learning for Adaptive Multimodal Representation in the Taobao Shangou Recommender System

- 三阶段设计:先预训练,再用行为反馈优化嵌入,最后自适应融合
- 线上测试提升0.55%订单量,离线AUC增0.12~0.20,PCOC更优
- 适合工业级推荐系统中多模态对齐与用户意图动态建模场景
现代外卖推荐系统日益依赖图像、文本和用户行为等多模态信号以提升体验,但异构模态的有效融合仍具挑战,阻碍了多模态联合建模与用户意图的动态适应。主流两阶段方法中,图像-文本编码器的内容语义预训练与行为驱动排序模型分离,导致语义理解与用户行为模式之间对齐不足。为此,我们提出GALA,一个三阶段流程,核心创新在于中间的“生成式强化学习对齐”阶段:基于用户行为生成多模态预训练数据,并通过基于转换的奖励进行优化,有效弥合预训练与微调之间的差距。GALA包含三阶段:第一阶段,在搜索日志的查询-图像-文本三元组上进行行为感知的三元组预训练,早期捕捉用户意图与内容偏好;第二阶段,引入新颖的奖励驱动优化(GRPO)阶段,动态对齐多模态嵌入与用户行为,填补预训练-微调鸿沟;第三阶段,通过自适应门控融合多模态与ID嵌入,并采用混合损失,确保长期以ID为主导的训练中仍保留多模态贡献。GALA已在淘宝闪购生产环境部署,服务超2亿日活用户。相比最先进方法,离线表现稳定提升+0.12/+0.20 AUC,PCOC指标更优。大规模线上A/B测试进一步验证其有效性,订单量提升0.55%,证明其在工业规模下的卓越性能与跨需求场景的鲁棒性。
原文摘要 · Abstract (English)
Modern recommender systems in food delivery increasingly leverage multimodal signals, including images, text, and user interaction histories, to enhance user experience, yet effective fusion of these heterogeneous modalities remains challenging, hindering both the joint modeling of multimodal signals and adaptation to evolving user intent. In mainstream two-stage approaches, the separation between content-semantic pretraining of image-text encoders and behavior-driven ranking models limits alignment between semantic understanding and user behavior patterns. To address these issues, we present GALA, a three-stage pipeline whose core innovation lies in an intermediate "generative RL alignment" stage that constructs multimodal pretraining data from user behavior and refines it via conversion-based rewards, effectively bridging the pretraining-fine-tuning gap to align with downstream objectives. GALA comprises three stages: first, behavior-aware triplet pretraining on query-image-text pairs from search logs to early capture user intent and content preferences; second, a novel intermediate stage that refines multimodal embeddings through reward-driven optimization (GRPO) to dynamically align them with user behavior and bridge the pretraining-fine-tuning gap; and finally, integration of multimodal and ID embeddings via adaptive gating with a hybrid loss, preserving multimodal contributions under long-term ID-dominant training. GALA has been deployed in the production environment at Taobao Shangou, serving over 200 million daily active users. Compared with state-of-the-art (SOTA) methods, it delivers consistent offline gains of +0.12/+0.20 AUC along with better PCOC metrics. Large-scale online A/B tests further report a 0.55 percent increase in order volume, confirming GALA's effectiveness at industrial scale and its robustness across diverse demand patterns.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。