用大模型生成商业意图,实时高效召回广告
Real-time Ad retrieval via LLM-generative Commercial Intention for Sponsored Search Advertising

- 用大模型生成商业意图文本作为中间表示,替代传统ID映射
- 线上实测带来5.04%消费增长、6.37%GMV提升和1.28%点击率改善
- 适合大规模广告系统,尤其看重实时性与可扩展性的场景
将大语言模型(LLM)与检索系统结合,在广告检索中展现出巨大潜力。现有方法通过生成数值或内容型文档ID来召回广告,但存在数值ID与文档的单对多映射问题,且内容提取耗时,导致语义效率低,难以在大规模语料中扩展。本文提出实时广告检索框架RARE,利用定制化大模型生成的商业意图(CIs)作为中间语义表征,直接实现查询到广告的实时检索。这些商业意图由注入商业知识的模型生成,提升了领域相关性,且每个意图对应多个广告,形成轻量且可扩展的集合。RARE已在真实在线系统中部署,日均处理搜索量达数亿级。线上实验显示:消费提升5.04%,GMV增长6.37%,点击率提高1.28%,浅层转化率上升5.29%。离线实验表明,RARE在四大类共十种基线方法中表现更优。
原文摘要 · Abstract (English)
The integration of Large Language Models (LLMs) with retrieval systems has shown promising potential in retrieving documents (docs) or advertisements (ads) for a given query. Existing LLM-based retrieval methods generate numeric or content-based DocIDs to retrieve docs/ads. However, the one-to-few mapping between numeric IDs and docs, along with the time-consuming content extraction, leads to semantic inefficiency and limits scalability in large-scale corpora. In this paper, we propose the Real-time Ad REtrieval (RARE) framework, which leverages LLM-generated text called Commercial Intentions (CIs) as an intermediate semantic representation to directly retrieve ads for queries in real-time. These CIs are generated by a customized LLM injected with commercial knowledge, enhancing its domain relevance. Each CI corresponds to multiple ads, yielding a lightweight and scalable set of CIs. RARE has been implemented in a real-world online system, handling daily search volumes in the hundreds of millions. The online implementation has yielded significant benefits: a 5.04% increase in consumption, a 6.37% rise in Gross Merchandise Volume (GMV), a 1.28% enhancement in click-through rate (CTR) and a 5.29% increase in shallow conversions. Extensive offline experiments show RARE's superiority over ten competitive baselines in four major categories.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。