提出新指标JaccDiv,量化音乐营销文本生成多样性。
JaccDiv: A Metric and Benchmark for Quantifying Diversity of Generated Marketing Text in the Music Industry
- 用T5、GPT-3.5等模型结合微调/少样本/零样本生成文本。
- 引入JaccDiv指标,可有效评估生成文本的多样性水平。
- 适合关注自动化内容多样性的营销与传媒领域研究者。
在线平台日益依赖数据到文本技术自动生成内容以辅助用户。然而,传统生成方法常陷入重复模式,仅经过几次迭代后便产生单调乏味的文本集合。本文研究基于大语言模型的数据到文本方法,旨在自动生成质量高且多样性足够的音乐行业营销文本。我们采用T5、GPT-3.5、GPT-4和LLaMa2等语言模型,结合微调、少样本及零样本策略建立基线。同时,提出全新指标JaccDiv,用于评估一组文本的多样性。本研究不仅适用于音乐产业,对其他存在自动化内容重复问题的领域也具广泛参考价值。
原文摘要 · Abstract (English)
Online platforms are increasingly interested in using Data-to-Text technologies to generate content and help their users. Unfortunately, traditional generative methods often fall into repetitive patterns, resulting in monotonous galleries of texts after only a few iterations. In this paper, we investigate LLM-based data-to-text approaches to automatically generate marketing texts that are of sufficient quality and diverse enough for broad adoption. We leverage Language Models such as T5, GPT-3.5, GPT-4, and LLaMa2 in conjunction with fine-tuning, few-shot, and zero-shot approaches to set a baseline for diverse marketing texts. We also introduce a metric JaccDiv to evaluate the diversity of a set of texts. This research extends its relevance beyond the music industry, proving beneficial in various fields where repetitive automated content generation is prevalent.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。