用人类反馈训练小模型讲故事,效率比传统方法高400倍
Once Upon a Time: Interactive Learning for Storytelling with Small Language Models
- 让小模型生成故事,由大模型从可读性等维度给出高阶反馈
- 仅需100万词交互学习,效果相当于4.1亿词纯文本训练
- 适合资源有限但需高质量生成任务的研究者与开发者
儿童习得语言不仅依赖听觉输入,更通过社会互动。相比之下,大型语言模型通常基于海量文本进行下一步词预测训练。受此启发,我们探索是否可通过引入高层认知反馈,使语言模型在更少数据下提升能力。我们训练一个学生模型生成故事,由教师模型从可读性、叙事连贯性和创造性三方面评分。通过调整预训练数据量,评估这种交互式学习对形式与功能语言能力的影响。结果表明,高层反馈具有极高的数据效率:仅需100万词的交互学习,讲故事能力提升幅度相当于4.1亿词的下一步词预测训练。
原文摘要 · Abstract (English)
Children efficiently acquire language not just by listening, but by interacting with others in their social environment. Conversely, large language models are typically trained with next-word prediction on massive amounts of text. Motivated by this contrast, we investigate whether language models can be trained with less data by learning not only from next-word prediction but also from high-level, cognitively inspired feedback. We train a student model to generate stories, which a teacher model rates on readability, narrative coherence, and creativity. By varying the amount of pretraining before the feedback loop, we assess the impact of this interactive learning on formal and functional linguistic competence. We find that the high-level feedback is highly data efficient: With just 1 M words of input in interactive learning, storytelling skills can improve as much as with 410 M words of next-word prediction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。