构建大规模广告文案改写数据集,助力生成更吸引人的广告内容
AdParaphrase v2.0: Generating Attractive Ad Texts Using a Preference-Annotated Paraphrase Dataset
- 构建16,460对广告文案改写样本,每对经10人评分,标注偏好数据
- 发现新语言特征,提升广告吸引力,验证偏好与实际表现的相关性
- 提供无需参考的LLM评估方法,适合广告生成与优化研究者使用
识别使广告文案更具吸引力的因素对广告成功至关重要。本文提出AdParaphrase v2.0,一个包含人类偏好标注的广告文案改写数据集,用于分析语言特征并支持吸引人广告文案生成方法的研发。相较于v1.0,该数据集规模扩大20倍,包含16,460组广告文案改写对,每对由十名评估者标注偏好数据,实现更全面可靠的分析。实验揭示了v1.0中未观察到的多个吸引人广告文案的语言特征,并探索了多种生成策略。此外,分析表明人类偏好与广告表现存在关联,并展示了基于大语言模型的无参考指标在评估广告吸引力方面的潜力。数据集已公开:https://github.com/CyberAgentAILab/AdParaphrase-v2.0。
原文摘要 · Abstract (English)
Identifying factors that make ad text attractive is essential for advertising success. This study proposes AdParaphrase v2.0, a dataset for ad text paraphrasing, containing human preference data, to enable the analysis of the linguistic factors and to support the development of methods for generating attractive ad texts. Compared with v1.0, this dataset is 20 times larger, comprising 16,460 ad text paraphrase pairs, each annotated with preference data from ten evaluators, thereby enabling a more comprehensive and reliable analysis. Through the experiments, we identified multiple linguistic features of engaging ad texts that were not observed in v1.0 and explored various methods for generating attractive ad texts. Furthermore, our analysis demonstrated the relationships between human preference and ad performance, and highlighted the potential of reference-free metrics based on large language models for evaluating ad text attractiveness. The dataset is publicly available at: https://github.com/CyberAgentAILab/AdParaphrase-v2.0.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。