用大模型理解广告主意图,减少误判,提升广告安全审查体验。
Advertiser Content Understanding via LLMs for Google Ads Safety
- 基于广告主多维度数据构建内容画像,输入大模型判断其违规可能性。
- 微调后在小样本测试集上达到95%准确率,有效降低误判。
- 适合关注广告安全与用户体验优化的平台方及算法团队参考。
Google Ads 的广告内容安全需对数十亿条广告进行内容政策分类。一致且准确的政策执行对广告主体验和用户安全至关重要,但实现难度高,因此提升该能力对广告主和用户均有显著价值。政策执行不一致会增加合规摩擦,导致优质广告主体验受损,而劣质广告主则通过批量生成相似广告,利用漏洞绕过审查。本研究提出一种基于大语言模型(LLMs)的方法,用于理解广告主意图以识别内容政策违规行为。重点在于识别优质广告主,减少内容误判,改善广告主体验,该方法也可轻松扩展至识别劣质广告主。通过整合广告主广告、域名、定向信息等多源信号生成内容画像,并结合大模型对广告主、产品或品牌的已有知识,判断其是否可能违反特定政策。经少量提示工程微调后,该方法在小规模测试集上达到95%准确率。
原文摘要 · Abstract (English)
Ads Content Safety at Google requires classifying billions of ads for Google Ads content policies. Consistent and accurate policy enforcement is important for advertiser experience and user safety and it is a challenging problem, so there is a lot of value for improving it for advertisers and users. Inconsistent policy enforcement causes increased policy friction and poor experience with good advertisers, and bad advertisers exploit the inconsistency by creating multiple similar ads in the hope that some will get through our defenses. This study proposes a method to understand advertiser's intent for content policy violations, using Large Language Models (LLMs). We focus on identifying good advertisers to reduce content over-flagging and improve advertiser experience, though the approach can easily be extended to classify bad advertisers too. We generate advertiser's content profile based on multiple signals from their ads, domains, targeting info, etc. We then use LLMs to classify the advertiser content profile, along with relying on any knowledge the LLM has of the advertiser, their products or brand, to understand whether they are likely to violate a certain policy or not. After minimal prompt tuning our method was able to reach 95\% accuracy on a small test set.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。