对比细粒度人格提示生成数据的词汇多样性,发现细节增益有限。
Measuring Lexical Diversity of Synthetic Data Generated through Fine-Grained Persona Prompting
- 用多种词汇多样性指标量化生成文本的差异性
- 细粒度人格提示对多样性提升效果不显著
- 大模型下人格提示优于无提示,但不如长度限制有效
近期,细粒度人格提示被用于生成用于预训练和监督微调大语言模型的‘多样化’合成数据。本文通过一系列词汇多样性和冗余度指标,测量了基于人格提示生成的指令与回复的多样性。结果表明,合成指令/回复的多样性显著低于人类撰写内容。我们进一步采样不同规模的语言模型,在细粒度与粗粒度人格描述下生成回复,探究人格描述中细粒度信息对生成文本多样性的影响。结果显示,相比无人格提示,人格提示能提升词汇多样性,尤其在大模型中更明显;然而,增加细粒度人格细节带来的多样性提升极为有限,甚至不及仅在提示中添加长度限制的效果。
原文摘要 · Abstract (English)
Fine-grained personas have recently been used for generating 'diverse' synthetic data for pre-training and supervised fine-tuning of Large Language Models (LLMs). In this work, we measure the diversity of persona-driven synthetically generated prompts and responses with a suite of lexical diversity and redundancy metrics. First, we find that synthetic prompts/instructions are significantly less diverse than human-written ones. Next, we sample responses from LLMs of different sizes with fine-grained and coarse persona descriptions to investigate how much fine-grained detail in persona descriptions contribute to generated text diversity. Our results indicate that persona prompting produces higher lexical diversity than prompting without personas, particularly in larger models. In contrast, adding fine-grained persona details yields minimal gains in diversity compared to simply specifying a length cutoff in the prompt.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。