arXiv:2409.11547cs.CLcs.AI2024-09中稿 · as Main Conference…被引 39

小模型在短篇创作中表现超越人类,展现惊人创造力。

Small Language Models can Outperform Humans in Short Creative Writing: A Study Comparing SLMs with Humans and LLMs

  • 用微调的BART-large生成故事,比普通人更流畅有吸引力。
  • 小模型故事中15%含意外关联,远超大模型的3%。
  • 适合关注模型效率与创意平衡的研究者或创作者。

本文评估了微调后的小语言模型(SLM)BART-large在创意小说写作中的表现,并与人类作者及两个大语言模型(LLM)GPT-3.5和GPT-4o进行对比。实验包括两项:(i) 68名参与者对人类与SLM生成的故事在语法、相关性、创意和吸引力方面打分;(ii) 对各模型生成文本的语言特征进行定性分析。结果显示,BART-large整体得分2.11,高于人类平均分1.85,相对提升14%,尽管人类在创意上略有优势但不显著。定性分析表明,虽然GPT-4o语义连贯性近乎完美且陈词较少,但其语言更可预测,仅3%的摘要包含意外关联,而BART-large达15%。研究揭示模型规模与微调如何影响创作中创意、流畅性与连贯性的权衡,证明小模型在特定情境下可媲美甚至超越人类与大模型。

原文摘要 · Abstract (English)

In this paper, we evaluate the creative fiction writing abilities of a fine-tuned small language model (SLM), BART-large, and compare its performance to human writers and two large language models (LLMs): GPT-3.5 and GPT-4o. Our evaluation consists of two experiments: (i) a human study in which 68 participants rated short stories from humans and the SLM on grammaticality, relevance, creativity, and attractiveness, and (ii) a qualitative linguistic analysis examining the textual characteristics of stories produced by each model. In the first experiment, BART-large outscored average human writers overall (2.11 vs. 1.85), a 14% relative improvement, though the slight human advantage in creativity was not statistically significant. In the second experiment, qualitative analysis showed that while GPT-4o demonstrated near-perfect coherence and used less cliche phrases, it tended to produce more predictable language, with only 3% of its synopses featuring surprising associations (compared to 15% for BART). These findings highlight how model size and fine-tuning influence the balance between creativity, fluency, and coherence in creative writing tasks, and demonstrate that smaller models can, in certain contexts, rival both humans and larger models.

小模型创意写作生成质量模型对比

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。