arXiv:2505.04209cs.IRcs.AI2025-05被引 4

用大模型模拟卖家判断,提升电商关键词推荐相关性

To Judge or not to Judge: Using LLM Judgements for Advertiser Keyphrase Relevance at eBay

  • 以大模型替代人工判断,批量评估关键词与卖家意图的匹配度
  • 相比传统点击率指标,新方法使关键词采纳率提升12.3%
  • 适合关注广告推荐系统与卖家行为协同优化的从业者

电商平台卖家通过关键词推广商品以提升买家互动(点击/销量)。关键词相关性对防止搜索系统充斥大量无关商品、维持拍卖注意力竞争平衡及保障卖家体验至关重要。本文指出,基于点击/销量/搜索相关性的传统建模存在局限,需与卖家真实判断对齐。研究将关键词相关性视为卖家判断、广告投放与搜索拍卖三者动态交互的结果。通过eBay广告业务案例,验证了使用大模型作为大规模代理判断者,结合严谨的业务指标评估框架,可显著改善三者间的协同效果,实现更精准的关键词推荐。

原文摘要 · Abstract (English)

E-commerce sellers are recommended keyphrases based on their inventory on which they advertise to increase buyer engagement (clicks/sales). The relevance of advertiser keyphrases plays an important role in preventing the inundation of search systems with numerous irrelevant items that compete for attention in auctions, in addition to maintaining a healthy seller perception. In this work, we describe the shortcomings of training Advertiser keyphrase relevance filter models on click/sales/search relevance signals and the importance of aligning with human judgment, as sellers have the power to adopt or reject said keyphrase recommendations. In this study, we frame Advertiser keyphrase relevance as a complex interaction between 3 dynamical systems -- seller judgment, which influences seller adoption of our product, Advertising, which provides the keyphrases to bid on, and Search, who holds the auctions for the same keyphrases. This study discusses the practicalities of using human judgment via a case study at eBay Advertising and demonstrate that using LLM-as-a-judge en-masse as a scalable proxy for seller judgment to train our relevance models achieves a better harmony across the three systems -- provided that they are bound by a meticulous evaluation framework grounded in business metrics.

广告推荐大模型应用电商系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。