首个专用于生成式营销广告注入的评测基准,解决广告与体验平衡难题。
GEM-Bench: A Benchmark for Ad-Injected Response Generation within Generative Engine Marketing
- 构建三类数据集覆盖对话与搜索场景,支持多维度评估。
- 提示工程法提升点击率但降低用户满意度,预生成法缓解此问题但增加开销。
- 适合研究生成式广告、用户体验优化及商业化模型的开发者与学者。
生成式引擎营销(GEM)是一种新兴生态,通过在生成式引擎(如基于大语言模型的聊天机器人)响应中无缝嵌入相关广告来实现变现。其核心在于广告注入响应的生成与评估。然而,现有评测基准未专门针对此目标设计,制约了后续研究。为此,我们提出 GEM-Bench,首个面向 GEM 中广告注入响应生成的综合性评测基准。该基准包含三个精心构建的数据集,覆盖聊天与搜索场景;一套多维用户满意度与参与度度量体系;以及在可扩展多智能体框架中实现的多个基线方案。初步结果显示,简单提示方法虽能实现合理点击率,但常损害用户满意度;而基于预生成无广告响应插入广告的方法可缓解此问题,但引入额外计算开销。这些发现凸显了未来在 GEM 中设计更高效、更有效广告注入方案的必要性。评测基准及所有资源已公开:https://gem-bench.org/。
原文摘要 · Abstract (English)
Generative Engine Marketing (GEM) is an emerging ecosystem for monetizing generative engines, such as LLM-based chatbots, by seamlessly integrating relevant advertisements into their responses. At the core of GEM lies the generation and evaluation of ad-injected responses. However, existing benchmarks are not specifically designed for this purpose, which limits future research. To address this gap, we propose GEM-Bench, the first comprehensive benchmark for ad-injected response generation in GEM. GEM-Bench includes three curated datasets covering both chatbot and search scenarios, a metric ontology that captures multiple dimensions of user satisfaction and engagement, and several baseline solutions implemented within an extensible multi-agent framework. Our preliminary results indicate that, while simple prompt-based methods achieve reasonable engagement such as click-through rate, they often reduce user satisfaction. In contrast, approaches that insert ads based on pre-generated ad-free responses help mitigate this issue but introduce additional overhead. These findings highlight the need for future research on designing more effective and efficient solutions for generating ad-injected responses in GEM. The benchmark and all related resources are publicly available at https://gem-bench.org/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。