用因果优化提升电商搜索长期用户价值预测效果
DCEO: Direct Causal Effect Optimization for Long-Term User Value Modeling in E-commerce Search
- 基于因果效应直接优化用户级目标的融合权重
- 线上实验提升GMV 0.36%,优于传统代理指标
- 适合追求长期用户价值的电商平台搜索系统
工业级电商搜索系统最终目标是优化用户层面的长期指标,如用户n日累计购买或商品交易总额(GMV)。然而,搜索排序基于请求内的物品级评分,而长期目标在用户层面定义。现有方法通常通过手动设计的多目标融合方案来弥合粒度差异,将点击、加购、购买等物品级目标预测结果组合成排序分作为最终目标的代理。这类手工融合依赖少量可调权重,难以实现细粒度个性化,且与最终目标对齐不佳。本文提出DCEO(Direct Causal Effect Optimization),一种数据驱动框架,用于学习更贴近最终目标的物品级代理评分。首先将物品级代理评分聚合为用户级代理指标,并通过相对因果效应量化其与最终目标的对齐程度。随后构建演员-评论家框架:评论家估计给定用户级代理指标下的最终目标值,演员动态生成上下文相关的多目标融合权重,以构建物品级代理评分,并直接优化相对因果效应。大量离线实验和分析验证了DCEO的有效性与可解释性。此外,DCEO已在大规模工业电商搜索系统中部署,在41天线上A/B测试中,相比传统GMV代理提升了0.36%的GMV。
原文摘要 · Abstract (English)
Industrial e-commerce search systems ultimately aim to optimize the user-level long-term objective, such as n-day cumulative purchases or gross merchandise value (GMV) per user. However, such objectives are defined at the user level, whereas search ranking is based on item-level scores within each request. Existing methods typically bridge this granularity gap through manually designed multi-objective fusion, where predictions of multiple item-level objectives, such as clicks, carts, purchases, and transaction value, are combined into a ranking score that serves as a proxy for the ultimate objective. Such hand-crafted fusion schemes rely on a small set of manually tuned weights, limiting fine-grained personalization and leading to suboptimal alignment with the ultimate objective. In this paper, we propose DCEO (Direct Causal Effect Optimization), a data-driven framework for learning item-level proxy scores that are better aligned with the ultimate objective. We first aggregate the item-level proxy scores into a user-level proxy metric and quantify its alignment with the ultimate objective using a relative causal effect. We then develop an actor-critic framework, where the critic estimates the ultimate objective for a given user-level proxy metric, and the actor dynamically generates context-dependent fusion weights over multiple objectives to construct the item-level proxy scores and is trained to directly optimize the relative causal effect. Extensive offline experiments and analyses demonstrate the effectiveness and interpretability of DCEO. In addition, DCEO has been deployed in a large-scale industrial e-commerce search system, outperforming the conventional GMV proxy by 0.36% in GMV in a 41-day online A/B test.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。