通过控制提示约束数量,自动评估大模型的创造力。
CS4: Measuring the Creativity of Large Language Models Automatically by Controlling the Number of Story-Writing Constraints
- 用不同数量的约束条件设计提示,逐步提升创作难度。
- 模型在高约束下生成故事与训练数据重合度显著上升。
- 适合研究模型创意能力与指令遵循平衡的研究者使用。
评估大语言模型在故事创作中的创造力十分困难,因为模型生成的故事可能看似创新,实则与训练语料库中大量现有故事高度相似。为解决此问题,我们提出一个新型基准数据集CS4(通过控制合成约束特异性来比较创作技能),通过增加提示中的要求/约束数量,提升提示的明确性,从而阻碍模型复述训练数据中的高质量叙事。由此,无需人工标注即可间接衡量模型的创造力。我们在LLaMA、Gemma和Mistral上的实验揭示了模型在高约束提示下面临严重创意挑战,且不同模型在不同约束数量下的表现差异显著,展现出指令遵循能力与叙事连贯性之间的不同平衡。此外,对OLMo的实验表明,基于人类反馈的学习(LHF)虽能帮助模型从训练数据中选出更优故事,但对生成训练数据中未见的新颖故事能力提升有限。该基准已开源:https://github.com/anirudhlakkaraju/cs4_benchmark。
原文摘要 · Abstract (English)
Evaluating the creativity of large language models (LLMs) in story writing is difficult because LLM-generated stories could seemingly look creative but be very similar to some existing stories in their huge and proprietary training corpus. To overcome this challenge, we introduce a novel benchmark dataset with varying levels of prompt specificity: CS4 ($\mathbf{C}$omparing the $\mathbf{S}$kill of $\mathbf{C}$reating $\mathbf{S}$tories by $\mathbf{C}$ontrolling the $\mathbf{S}$ynthesized $\mathbf{C}$onstraint $\mathbf{S}$pecificity). By increasing the number of requirements/constraints in the prompt, we can increase the prompt specificity and hinder LLMs from retelling high-quality narratives in their training data. Consequently, CS4 empowers us to indirectly measure the LLMs' creativity without human annotations. Our experiments on LLaMA, Gemma, and Mistral not only highlight the creativity challenges LLMs face when dealing with highly specific prompts but also reveal that different LLMs perform very differently under different numbers of constraints and achieve different balances between the model's instruction-following ability and narrative coherence. Additionally, our experiments on OLMo suggest that Learning from Human Feedback (LHF) can help LLMs select better stories from their training data but has limited influence in boosting LLMs' ability to produce creative stories that are unseen in the training corpora. The benchmark is released at https://github.com/anirudhlakkaraju/cs4_benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。