少即是多:用极少数据也能让大模型学会复杂数学推理
LIMO: Less is More for Reasoning
- 仅用1%数据通过简单微调,激活模型推理能力
- AIME24准确率63.3%,超越此前模型6.5%的纪录
- 适合关注高效训练与认知模板设计的研究者
我们挑战了大语言模型复杂推理需海量训练数据的普遍认知。通过简单监督微调,我们的模型LIMO仅用1%的数据便在AIME24上达到63.3%准确率,在MATH500上达95.6%,显著优于此前微调模型(分别为6.5%和59.2%)。同时,LIMO在跨分布任务中表现出强劲泛化能力,绝对提升达45.8%,超越使用100倍数据训练的模型。由此我们提出‘少即是多推理假说’:在预训练已充分编码领域知识的基座模型中,复杂推理可通过少量精心设计的认知过程示范激发。该假说认为,推理涌现的关键不在于任务难度,而在于(1)模型预训练知识的完备性,以及(2)后训练示例作为‘认知模板’的有效引导。
原文摘要 · Abstract (English)
We challenge the prevailing assumption that complex reasoning in large language models (LLMs) necessitates massive training data. We demonstrate that sophisticated mathematical reasoning can emerge with only a few examples. Specifically, through simple supervised fine-tuning, our model, LIMO, achieves 63.3\% accuracy on AIME24 and 95.6\% on MATH500, surpassing previous fine-tuned models (6.5\% on AIME24, 59.2\% on MATH500) while using only 1\% of the training data required by prior approaches. Furthermore, LIMO exhibits strong out-of-distribution generalization, achieving a 45.8\% absolute improvement across diverse benchmarks, outperforming models trained on 100x more data. Synthesizing these findings, we propose the Less-Is-More Reasoning Hypothesis (LIMO Hypothesis): In foundation models where domain knowledge has been comprehensively encoded during pre-training, sophisticated reasoning can emerge through minimal but strategically designed demonstrations of cognitive processes. This hypothesis suggests that the threshold for eliciting complex reasoning is not dictated by task complexity but rather by two key factors: (1) the completeness of the model's pre-trained knowledge base and (2) the effectiveness of post-training examples in serving as "cognitive templates" that guide reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。