用大模型生成精准广告文案并自动评估,提升电商外投效果
LLMs for Customized Marketing Content Generation and Evaluation at Scale
- 融合多源数据生成关键词定制广告文案,减少人工干预
- 实测点击率高9%,曝光量增12%,每千次展示成本降0.38%
- 自动化评估系统与人类协作,降低审核成本,保持质量一致性
外部营销在电商中至关重要,帮助商家通过外部平台触达用户并引流至官网。然而当前多数营销内容过于通用、模板化,且与落地页匹配度低,影响转化效果。为此,我们提出MarketingFM——一种检索增强型系统,整合多源数据生成关键词相关的广告文案,实现极低人工参与。通过离线人工与自动化评估及大规模在线A/B测试验证,结果显示,关键词导向的广告文案相比模板提升最高9%的点击率(CTR),增加12%的曝光量,同时将每千次展示成本(CPC)降低0.38%,显著改善广告排名与成本效率。尽管如此,人工审核生成内容仍耗时费力。为此我们提出AutoEval-Main,结合规则指标与大模型作为裁判(LLM-as-a-Judge)技术,确保文案符合营销准则。在大规模人工标注实验中,该系统与人工评审者达成89.57%的一致性。进一步提出AutoEval-Update,一种低成本的LLM-人类协同框架,通过有选择地采样代表性广告进行人工审查,并由批判性大模型生成对齐报告,动态优化评估提示,适应变化标准。实验表明,该机制可提出有意义的改进建议,提高大模型与人工评价一致性。但最终阈值设定与改进验证仍需人类把关。
原文摘要 · Abstract (English)
Offsite marketing is essential in e-commerce, enabling businesses to reach customers through external platforms and drive traffic to retail websites. However, most current offsite marketing content is overly generic, template-based, and poorly aligned with landing pages, limiting its effectiveness. To address these limitations, we propose MarketingFM, a retrieval-augmented system that integrates multiple data sources to generate keyword-specific ad copy with minimal human intervention. We validate MarketingFM via offline human and automated evaluations and large-scale online A/B tests. In one experiment, keyword-focused ad copy outperformed templates, achieving up to 9% higher CTR, 12% more impressions, and 0.38% lower CPC, demonstrating gains in ad ranking and cost efficiency. Despite these gains, human review of generated ads remains costly. To address this, we propose AutoEval-Main, an automated evaluation system that combines rule-based metrics with LLM-as-a-Judge techniques to ensure alignment with marketing principles. In experiments with large-scale human annotations, AutoEval-Main achieved 89.57% agreement with human reviewers. Building on this, we propose AutoEval-Update, a cost-efficient LLM-human collaborative framework to dynamically refine evaluation prompts and adapt to shifting criteria with minimal human input. By selectively sampling representative ads for human review and using a critic LLM to generate alignment reports, AutoEval-Update improves evaluation consistency while reducing manual effort. Experiments show the critic LLM suggests meaningful refinements, improving LLM-human agreement. Nonetheless, human oversight remains essential for setting thresholds and validating refinements before deployment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。