arXiv:2502.11569cs.CLcs.AI2025-02中稿 · EMNLP被引 30

小模型也能强推理,关键在训练方法而非规模。

Towards Reasoning Ability of Small Language Models

  • 构建首个系统评估小模型推理能力的基准ThinkSLM。
  • 量化保留推理能力,剪枝严重破坏性能,数据质量影响更大。
  • 部分小模型表现媲美甚至超越大模型,适合资源受限场景。

推理曾被视为大语言模型的涌现特性,但近期研究挑战这一观点,表明小语言模型(SLMs)也能达到竞争性推理水平。本文提出ThinkSLM,首个系统评估从零训练或通过量化、剪枝、蒸馏衍生的小模型推理能力的基准。我们建立可靠评估标准,对比现有方法与人类评估结果。对6个主流模型家族的72个SLMs在17个推理基准上进行评估,实验重复3次以确保稳健性。结果表明:1)小模型的推理能力受训练方法和数据质量显著影响,而非仅依赖模型规模;2)量化能较好保持推理能力,而剪枝会显著破坏性能;3)大模型对对抗扰动和中间推理更鲁棒,但某些小模型表现接近甚至超过大模型。研究质疑了‘规模唯一’的推理提升路径,预示未来可通过结构化训练或后处理压缩开发具备强推理能力的小模型。ThinkSLM排行榜已公开:https://ctrl-gaurav.github.io/thinkslm.github.io/

原文摘要 · Abstract (English)

Reasoning has long been viewed as an emergent property of large language models (LLMs). However, recent studies challenge this assumption, showing that small language models (SLMs) can also achieve competitive reasoning performance. This paper introduces ThinkSLM, the first extensive benchmark to systematically evaluate and study the reasoning abilities of SLMs trained from scratch or derived from LLMs through quantization, pruning, and distillation. We first establish a reliable evaluation criterion comparing available methods and LLM judges against our human evaluations. Then we present a study evaluating 72 diverse SLMs from six major model families across 17 reasoning benchmarks. We repeat all our experiments three times to ensure a robust assessment. Our findings show that: 1) reasoning ability in SLMs is strongly influenced by training methods and data quality rather than solely model scale; 2) quantization preserves reasoning capability, while pruning significantly disrupts it; 3) larger models consistently exhibit higher robustness against adversarial perturbations and intermediate reasoning, but certain smaller models closely match or exceed the larger models' performance. Our findings challenge the assumption that scaling is the only way to achieve strong reasoning. Instead, we foresee a future where SLMs with strong reasoning capabilities can be developed through structured training or post-training compression. Our ThinkSLM Leaderboard is publicly available at: https://ctrl-gaurav.github.io/thinkslm.github.io/

小模型推理能力量化模型压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。