用普通游戏显卡训练出高性能数学推理模型
Can A Gamer Train A Mathematical Reasoning Model?
- 结合强化学习与内存优化,在单张显卡上训练15亿参数模型
- 在3080 Ti上表现超越数倍更大的模型,基准测试成绩相当
- 让普通研究者也能低成本开展高级数学推理研究
尽管大语言模型在数学推理等任务中表现优异,但其训练通常需要高昂的计算资源。近期进展虽降低了成本,但仍依赖高端硬件集群。本文展示仅用一张普通游戏显卡(RTX 3080 Ti,16GB内存)即可训练出性能可靠的15亿参数数学推理模型,其在数学推理基准上的表现与更大模型相当甚至更优,适用于资源受限环境。该结果挑战了高性能数学推理必须依赖大规模基础设施的传统观念,推动高精度AI研究的普惠化。项目开源地址:https://github.com/shinandrew/YouronMath。
原文摘要 · Abstract (English)
While large language models (LLMs) have achieved remarkable performance in various tasks including mathematical reasoning, their development typically demands prohibitive computational resources. Recent advancements have reduced costs for training capable models, yet even these approaches rely on high-end hardware clusters. In this paper, we demonstrate that a single average gaming GPU can train a solid mathematical reasoning model, by integrating reinforcement learning and memory optimization techniques. Specifically, we train a 1.5B parameter mathematical reasoning model on RTX 3080 Ti of 16GB memory that achieves comparable or better performance on mathematical reasoning benchmarks than models several times larger, in resource-constrained environments. Our results challenge the paradigm that state-of-the-art mathematical reasoning necessitates massive infrastructure, democratizing access to high-performance AI research. https://github.com/shinandrew/YouronMath.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。