构建广告文案改写数据集,揭示吸引人的语言特征。
AdParaphrase: Paraphrase Dataset for Analyzing Linguistic Features toward Generating Attractive Ad Texts
- 通过语义等价但风格不同的广告文案对,分析人类偏好差异。
- 优选文案更流畅、更长、名词更多且使用括号符号。
- 基于发现优化生成模型,显著提升广告吸引力。
有效的语言选择在广告成功中起关键作用。本研究旨在探索影响人类偏好的广告文案语言特征。尽管吸引人的广告文案生成是热门研究方向,但对具体影响吸引力的语言特征的理解受限于多重障碍:其一,人类偏好复杂,受品牌名称等内容因素和语言风格等多重影响;其二,缺乏包含人类偏好信息的公开广告文本数据集,如广告表现指标或人工反馈。为此,我们提出 AdParaphrase,一个包含人类偏好标注的改写数据集,其中每对广告文案语义等价但措辞与风格不同。该数据集支持聚焦语言特征差异的偏好分析。我们的分析发现,被人类评委偏好的广告文案具有更高流畅性、更长长度、更多名词及括号符号的使用。此外,我们证明,结合这些发现的广告生成模型能显著提升文本吸引力。数据集已公开:https://github.com/CyberAgentAILab/AdParaphrase。
原文摘要 · Abstract (English)
Effective linguistic choices that attract potential customers play crucial roles in advertising success. This study aims to explore the linguistic features of ad texts that influence human preferences. Although the creation of attractive ad texts is an active area of research, progress in understanding the specific linguistic features that affect attractiveness is hindered by several obstacles. First, human preferences are complex and influenced by multiple factors, including their content, such as brand names, and their linguistic styles, making analysis challenging. Second, publicly available ad text datasets that include human preferences are lacking, such as ad performance metrics and human feedback, which reflect people's interests. To address these problems, we present AdParaphrase, a paraphrase dataset that contains human preferences for pairs of ad texts that are semantically equivalent but differ in terms of wording and style. This dataset allows for preference analysis that focuses on the differences in linguistic features. Our analysis revealed that ad texts preferred by human judges have higher fluency, longer length, more nouns, and use of bracket symbols. Furthermore, we demonstrate that an ad text-generation model that considers these findings significantly improves the attractiveness of a given text. The dataset is publicly available at: https://github.com/CyberAgentAILab/AdParaphrase.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。