arXiv:2601.02346cs.AI2026-01被引 11

7B小模型实现顶尖推理能力,高效且可扩展。

Falcon-H1R: Pushing the Reasoning Frontiers with a Hybrid Model for Efficient Test-Time Scaling

  • 采用混合并行架构与精准训练策略提升效率
  • 7B模型性能媲美甚至超越2-7倍大的模型
  • 适合需要长链推理和测试时扩展的场景

本文提出Falcon-H1R,一个70亿参数的推理优化小语言模型,证明了小型语言模型(SLMs)在推理任务上具备达到顶尖性能的可行性。Falcon-H1R在多种高推理强度基准测试中,表现持续优于或持平于规模大2至7倍的先进模型。这一成果凸显了精细数据筛选与针对性训练策略(包括高效SFT和强化学习缩放)的重要性,无需增大模型尺寸即可显著提升性能。此外,通过混合并行架构设计,Falcon-H1R在推理速度、生成效率和准确率方面实现了三重突破,进一步拓展了推理效率的三维边界。借助最新提出的DeepConf方法,该模型在测试时扩展效率上达到业界领先水平,在精度与计算成本间取得显著平衡。结果表明,通过针对性训练与架构设计,紧凑模型亦能实现强大且可扩展的推理能力。

原文摘要 · Abstract (English)

This work introduces Falcon-H1R, a 7B-parameter reasoning-optimized model that establishes the feasibility of achieving competitive reasoning performance with small language models (SLMs). Falcon-H1R stands out for its parameter efficiency, consistently matching or outperforming SOTA reasoning models that are $2\times$ to $7\times$ larger across a variety of reasoning-intensive benchmarks. These results underscore the importance of careful data curation and targeted training strategies (via both efficient SFT and RL scaling) in delivering significant performance gains without increasing model size. Furthermore, Falcon-H1R advances the 3D limits of reasoning efficiency by combining faster inference (through its hybrid-parallel architecture design), token efficiency, and higher accuracy. This unique blend makes Falcon-H1R-7B a practical backbone for scaling advanced reasoning systems, particularly in scenarios requiring extensive chain-of-thoughts generation and parallel test-time scaling. Leveraging the recently introduced DeepConf approach, Falcon-H1R achieves state-of-the-art test-time scaling efficiency, offering substantial improvements in both accuracy and computational cost. As a result, Falcon-H1R demonstrates that compact models, through targeted model training and architectural choices, can deliver robust and scalable reasoning performance.

推理模型小模型测试时扩展高效推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。