arXiv:2506.17863cs.CL2025-06被引 12

用大模型生成精准广告文案并自动评估,提升电商外投效果

LLMs for Customized Marketing Content Generation and Evaluation at Scale

  • 融合多源数据生成关键词定制广告文案,减少人工干预
  • 实测点击率高9%,曝光量增12%,每千次展示成本降0.38%
  • 自动化评估系统与人类协作,降低审核成本,保持质量一致性

外部营销在电商中至关重要,帮助商家通过外部平台触达用户并引流至官网。然而当前多数营销内容过于通用、模板化,且与落地页匹配度低,影响转化效果。为此,我们提出MarketingFM——一种检索增强型系统,整合多源数据生成关键词相关的广告文案,实现极低人工参与。通过离线人工与自动化评估及大规模在线A/B测试验证,结果显示,关键词导向的广告文案相比模板提升最高9%的点击率(CTR),增加12%的曝光量,同时将每千次展示成本(CPC)降低0.38%,显著改善广告排名与成本效率。尽管如此,人工审核生成内容仍耗时费力。为此我们提出AutoEval-Main,结合规则指标与大模型作为裁判(LLM-as-a-Judge)技术,确保文案符合营销准则。在大规模人工标注实验中,该系统与人工评审者达成89.57%的一致性。进一步提出AutoEval-Update,一种低成本的LLM-人类协同框架,通过有选择地采样代表性广告进行人工审查,并由批判性大模型生成对齐报告,动态优化评估提示,适应变化标准。实验表明,该机制可提出有意义的改进建议,提高大模型与人工评价一致性。但最终阈值设定与改进验证仍需人类把关。

原文摘要 · Abstract (English)

Offsite marketing is essential in e-commerce, enabling businesses to reach customers through external platforms and drive traffic to retail websites. However, most current offsite marketing content is overly generic, template-based, and poorly aligned with landing pages, limiting its effectiveness. To address these limitations, we propose MarketingFM, a retrieval-augmented system that integrates multiple data sources to generate keyword-specific ad copy with minimal human intervention. We validate MarketingFM via offline human and automated evaluations and large-scale online A/B tests. In one experiment, keyword-focused ad copy outperformed templates, achieving up to 9% higher CTR, 12% more impressions, and 0.38% lower CPC, demonstrating gains in ad ranking and cost efficiency. Despite these gains, human review of generated ads remains costly. To address this, we propose AutoEval-Main, an automated evaluation system that combines rule-based metrics with LLM-as-a-Judge techniques to ensure alignment with marketing principles. In experiments with large-scale human annotations, AutoEval-Main achieved 89.57% agreement with human reviewers. Building on this, we propose AutoEval-Update, a cost-efficient LLM-human collaborative framework to dynamically refine evaluation prompts and adapt to shifting criteria with minimal human input. By selectively sampling representative ads for human review and using a critic LLM to generate alignment reports, AutoEval-Update improves evaluation consistency while reducing manual effort. Experiments show the critic LLM suggests meaningful refinements, improving LLM-human agreement. Nonetheless, human oversight remains essential for setting thresholds and validating refinements before deployment.

广告生成大模型应用自动化评估电商营销

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。