arXiv:2507.15092cs.CL2025-07被引 1

提出新指标PATTR,解决生成文本长度差异导致的多样性评估偏差。

A Penalty Goes a Long Way: Measuring Lexical Diversity in Synthetic Texts Under Prompt-Influenced Length Variations

  • 引入任务目标长度$ L_T $,修正长度变化对词汇多样性评估的影响。
  • 在2000多万字视频脚本数据上验证,PATTR比传统指标更准确反映真实多样性。
  • 适合需要精准筛选高多样性生成内容的研究者与工程师使用。

大型语言模型生成的合成文本被广泛用于进一步训练和优化自身。词汇多样性对合成数据的有效性至关重要,研究者常通过提示工程提升多样性。然而,提示变化对响应长度的影响,以及由此带来的词汇多样性测量偏差,仍缺乏深入探讨。本文提出抗长度偏倚的类型-词项比率(PATTR)度量方法,利用来自LLaMA、OLMo和Phi系列共七种模型生成的超2000万词合成语料,聚焦于视频脚本创作这一需高多样性的任务。通过PATTR评估单条响应的词汇多样性,并与移动平均型类型-词项比率(MATTR)和压缩率(CR)对比。分析揭示:文本长度波动会系统性偏向较短回应。与现有方法不同,PATTR显式引入任务目标长度$ L_T $,有效缓解长度偏差。进一步实验表明,使用PATTR筛选前10/100/1,000条最多样化响应时,其在保持$ L_T $高符合度的前提下,多样性表现优于或相当甚至超越MATTR与CR。

原文摘要 · Abstract (English)

Synthetic text generated by Large Language Models (LLMs) is increasingly used for further training and improvement of LLMs. Diversity is crucial for the effectiveness of synthetic data, and researchers rely on prompt engineering to improve diversity. However, the impact of prompt variations on response text length, and, more importantly, the consequential effect on lexical diversity measurements, remain underexplored. In this work, we propose Penalty-Adjusted Type-Token Ratio (PATTR), a diversity metric robust to length variations. We generate a large synthetic corpus of over 20M words using seven models from the LLaMA, OLMo, and Phi families, focusing on a creative writing task of video script generation, where diversity is crucial. We evaluate per-response lexical diversity using PATTR and compare it against existing metrics of Moving-Average TTR (MATTR) and Compression Ratio (CR). Our analysis highlights how text length variations introduce biases favoring shorter responses. Unlike existing metrics, PATTR explicitly considers the task-specific target response length ($L_T$) to effectively mitigate length biases. We further demonstrate the utility of PATTR in filtering the top-10/100/1,000 most lexically diverse responses, showing that it consistently outperforms MATTR and CR by yielding on par or better diversity with high adherence to $L_T$.

文本生成多样性评估大模型提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。