小数据也能练出强推理能力,法语大模型效果显著提升
Pensez: Less Data, Better Reasoning -- Rethinking French LLM
- 仅用2000条精选双语数据做微调,提升模型推理与法语能力
- 在AIME25上准确率提升20%,法语数学题准确率提升12%
- 适合资源有限但需高效构建多语言强推理模型的场景
大型语言模型在自然语言处理任务中表现卓越,但在数学推理和非英语语言等专业领域常需海量数据训练。本文提出反向思路:通过在仅2000条精心筛选的英法双语数据上进行针对性监督微调(SFT),显著提升模型的推理能力与法语水平。实验表明,Pensez 7B模型在AIME25基准上准确率提升20%,在法语MATH level 5测试中提升12%。结果挑战了‘大规模数据是强推理能力前提’的普遍认知,凸显高质量数据策展与优化微调的潜力。研究为资源受限环境下高效构建高性能多语言大模型提供了新路径。
原文摘要 · Abstract (English)
Large language models (LLMs) have demonstrated remarkable capabilities in various natural language processing tasks. However, achieving strong performance in specialized domains like mathematical reasoning and non-English languages often requires extensive training on massive datasets. This paper investigates a contrasting approach: strategic fine-tuning on a small, high-quality, bilingual (English-French) dataset to enhance both the reasoning capabilities and French language proficiency of a large language model. Rather than relying on scale, we explore the hypothesis that targeted data curation and optimized training can achieve competitive, or even superior, performance. We demonstrate, through targeted supervised fine-tuning (SFT) on only 2,000 carefully selected samples, significant improvements in mathematical reasoning. Specifically, Pensez 7B exhibits an increase in accuracy of the base model up to 20% on the AIME25 and a 12% increase on a French MATH level 5 benchmark. These results challenge the prevailing assumption that massive datasets are aprerequisite for strong reasoning performance in LLMs, highlighting the potential of strategic data curation and optimized fine-tuning for enhancing both specialized skills and multilingual capabilities. Our findings have implications for the efficient development of high-performing, multilingual LLMs, especially in resource-constrained scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。