小模型通过数据学习国际象棋规则,能准确下棋解题。
Learning the Latent Rules of a Game from Data: A Chess Story
- 用1000到100万条指令微调2800万和1.25亿参数的小模型。
- 微调后模型能正确提出合法走法并解决棋局问题。
- 更多数据减少幻觉,多轮微调提升表现。
我们证明了参数量仅数百万的小型预训练生成语言模型能够从过程相关数据中学习其潜在规则。受斯蒂芬·茨威格小说《象棋的故事》启发,我们展示2800万和1.25亿参数的预训练小型语言模型(SLMs),在使用1000至100万条示例进行指令微调后,可学会国际象棋规则,提出合法走法,并准确解决棋局问题。我们还研究了连续微调轮次对结果的影响,发现增加指令微调样本数量可降低模型幻觉,提升性能。
原文摘要 · Abstract (English)
We demonstrate that small pretrained foundational generative language models with millions of parameters can learn the latent rules of a process from data associated with the process. Inspired by Stefan Zweig's novella "Schachnovelle," also known as "The Royal Game" in English, we show that 28M and 125M parameter pretrained foundational small language models (SLMs) can be instruction fine-tuned with 1,000-to-1,000,000 examples to learn the rules of chess, propose legal moves, and accurately solve chess problems. We also explore the impact of successive language model fine-tuning epochs on improved outcomes and demonstrate reductions in model hallucinations by increasing the number of instruction fine-tuning examples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。