检测大模型生成广告的风格变化,提升识别鲁棒性。
Detecting RAG Advertisements Across Advertising Styles
- 构建广告风格分类体系,涵盖显性与诉求类型维度。
- 发现轻量模型易受广告风格变化影响,检测效果下降。
- 基于实体识别的模型在多种风格下仍保持高精度。
大型语言模型(LLMs)使检索增强生成(RAG)系统能够生成融合上下文相关广告的自然响应,形成新型“生成式原生广告”。这种形式引发关注:能否自动检测此类广告?现有数据集未能反映营销文献中讨论的广告风格多样性。本文提出三点:(1) 构建面向LLM的广告风格分类体系,整合显性程度与诉求类型两个维度;(2) 模拟广告主通过改变风格规避检测的行为;(3) 评估多种广告检测方法在风格变化下的鲁棒性。扩展已有研究,我们训练了基于实体识别的模型,能精准定位广告内容,兼具高检测效果与强鲁棒性。考虑到广告拦截将在资源受限的终端设备上运行,我们纳入随机森林和SVM等轻量模型进行评估。结果显示,这些模型在风格变化下表现脆弱,凸显亟需高效且鲁棒的实用检测方案。
原文摘要 · Abstract (English)
Large language models (LLMs) enable a new form of advertising for retrieval-augmented generation (RAG) systems in which organic responses are blended with contextually relevant ads. The prospect of such "generated native ads" has sparked interest in whether they can be detected automatically. Existing datasets, however, do not reflect the diversity of advertising styles discussed in the marketing literature. In this paper, we (1) develop a taxonomy of advertising styles for LLMs, combining the style dimensions of explicitness and type of appeal, (2) simulate that advertisers may attempt to evade detection by changing their advertising style, and (3) evaluate a variety of ad-detection approaches with respect to their robustness under these changes. Expanding previous work on ad detection, we train models that use entity recognition to exactly locate an ad in an LLM response and find them to be both very effective at detecting responses with ads and largely robust to changes in the advertising style. Since ad blocking will be performed on low-resource end-user devices, we include lightweight models like random forests and SVMs in our evaluation. These models, however, are brittle under such changes, highlighting the need for further efficiency-oriented research for a practical approach to blocking of generated ads.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。