用微调大模型提升广告推荐的稳定性与可预测性。
LLM Retrieval for Stable and Predictable Ad Recommendations

- 用大模型提取广告创意的语义层级特征,构建语义感知候选生成
- 在大规模工业系统中显著提升稳定性和可预测性,同时改善点击率等指标
- 适合面临广告量激增与推荐可解释性挑战的系统
传统广告推荐系统主要关注点击或转化预测的准确性,采用召回率或归一化折现累积收益(NDCG)等标准指标。随着生成式AI技术推动广告库存和流动性急剧增长,预测的稳定性和可预测性日益关键。预测稳定性与可预测性可定义为系统对细微或噪声输入(广告、创意)扰动的鲁棒性,缺乏此类特性可能导致广告主感知的问题,如重复投放、冷启动和探索不足。本文提出一种新的评估框架,用于量化广告推荐系统的稳定性与可预测性,并展示一个在线验证的语义候选生成框架,该框架基于微调的大语言模型(LLM),通过根本性提升系统的语义感知能力,在多项指标上实现显著改进。该方法从广告创意中提取层次化语义属性,获得LLM表示,作为基于图的扩展基础,确保检索到的候选内容包含广告的语义变体,使广告主的小幅创意变动也能带来一致且可解释的用户投放结果。我们在大规模工业广告推荐系统中测试了该框架,在离线与在线A/B实验中均表现出显著提升,不仅增强可预测性,也改善了传统性能指标。尽管在广告系统中验证,该框架具有普适性,适用于任何面临类似扩展与可预测性挑战的大规模推荐与检索系统。
原文摘要 · Abstract (English)
Traditional ads recommendation systems have primarily focused on optimizing for prediction accuracy of click or conversion events using canonical metrics such as recall or normalized discounted cumulative gain (NDCG). With the hyper-growth of ads inventory and liquidity with generative AI technologies, the prediction stability and predictability is becoming increasingly critical. Intuitively, prediction stability and predictability can be defined to quantify system robustness with respect to minor/noisy input (ads, creatives) perturbations, the lack of which could lead to advertiser perceivable problems such as repeatability, cold start and under-exploration. In this paper, we introduce a new evaluation framework for quantifying stability and predictability of an ads recommender system, and present an online validated semantic candidate generation framework powered by fine-tuned Large Language Models (LLMs) that showed significant improvement along these metrics by fundamentally improving the semantic-awareness of the system. The approach extracts hierarchical semantic attributes from ad creatives to obtain LLM representations, which serve as the foundation for graph-based expansion, ensuring the retrieved candidates encapsulate semantic variants of an ad, guaranteeing that small creative variants from the advertiser yield consistent and explainable delivery results to the user. We tested this LLM ads retrieval framework in a large-scale industrial ads recommendation system, demonstrating significant improvements across offline and online A/B experiments, showcasing gains in both predictability and traditional performance metrics. Although evaluated in the ads stack, this is a general framework that can be applied broadly to any large-scale recommendation and retrieval systems facing similar scaling and predictability challenges.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。