用属性控制拓展创意写作数据,让模型学会写歌词、剧本等多样文体。
Scaling Creative Writing Beyond Story-Centric Data with Attribute-Guided Genre Expansion

- 用人类故事做种子,搭配人工标注的文体属性生成多样化创作。
- 构建了包含50,000条样本的13种文体语料库,覆盖说唱、歌词、剧本等。
- 模型在未见过的文体任务中表现更优,证明结构化扩展比单纯堆数据有效。
大型语言模型(LLMs)高质量的创意写作数据仍以故事为主,限制了模型对多种创意格式的结构与功能遵循能力。本文提出一种属性引导的文体扩展框架,实现创意写作数据的跨文体扩展。通过将主题广度与文体形式控制分离,利用人工撰写的故事情节作为多样创作种子,结合手工标注的文体属性来强制执行不同结构、风格和格式规范。由此生成强语言模型的符合文体要求的问答对,并经质量筛选后形成数据集。基于此框架,我们构建了涵盖13种创意文体的「Multi-Genre Collection」语料库,共50,000个样本,包括故事、说唱、歌词、剧本、游戏设计、角色设定等。在跨分布写作基准测试和保留文体诊断中,微调该数据集的模型显著优于基线模型及现有写作语料库训练模型。文体数量消融实验进一步表明,受控的文体扩展而非仅故事数据的规模扩大,是提升模型鲁棒创意写作能力的关键。
原文摘要 · Abstract (English)
High-quality creative writing data for large language models (LLMs) remains dominated by story-centric data, limiting models' ability to follow the structural and functional conventions of diverse creative formats. We propose an attribute-guided genre expansion framework for scaling creative writing data beyond story generation. By separating thematic breadth from genre-form control, our framework leverages human-authored story prompts as diverse creative seeds, while utilizing manually curated genre attributes to enforce distinct structural, stylistic, and formatting conventions. We combine these to prompt strong LLMs for genre-faithful query-response pairs, which are then quality-filtered. Applying this framework, we construct the Multi-Genre Collection, a 50K-example corpus spanning 13 creative genres, including story, rap, lyrics, scripts, game design, character design, and other creative formats. Experiments across out-of-distribution writing benchmarks and held-out genre diagnostics demonstrate that models fine-tuned on our data consistently surpass not only base models and writing-specialized baselines, but also models trained on existing writing corpora. Genre-count ablations further indicate that controlled genre expansion, rather than story-centric scaling alone, is a key driver of robust creative writing capability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。