用可调控的生成方法,打造两百万条中英文故事数据集。
Parameterized Synthetic Text Generation with SimpleStories
- 通过多层级参数化提示,实现对故事语法语义的规模化控制。
- 相比TinyStories,模型训练更高效且结果更易解释。
- 适合研究小参数语言模型与可控文本生成的学者使用。
我们提出SimpleStories,一个包含200万条样本的大型合成故事数据集,覆盖英语和日语。通过在多个抽象层次上参数化提示,实现了对故事特征的大规模可控生成,诱导出丰富的句法与语义多样性。针对新训练的模型套件进行消融实验表明,相比TinyStories数据集,该方法提升了样本效率并增强了模型可解释性。我们开源了模型构建的所有组成部分,旨在推动对端到端训练过程的深入研究。作为副产品,该工作将能生成语法正确自然语言的小参数语言模型的性能推向新前沿。
原文摘要 · Abstract (English)
We present SimpleStories, a large synthetic story dataset in simple language, consisting of 2 million samples each in English and Japanese. Through parameterizing prompts at multiple levels of abstraction, we achieve control over story characteristics at scale, inducing syntactic and semantic diversity. Ablations on a newly trained model suite show improved sample efficiency and model interpretability compared to the TinyStories dataset. We open-source all constituent parts of model creation, hoping to enable novel ways to study the end-to-end training process. As a byproduct, we move the frontier regarding the fewest-parameter language model that outputs grammatical natural language.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。