arXiv:2508.08636cs.CL2025-08被引 16

用1000+跨领域任务提升大模型推理能力,效果显著。

InternBootcamp: Boosting LLM Reasoning with Verifiable Task Scaling

  • 构建1000+跨领域任务环境,支持大模型推理训练。
  • 任务量提升两个数量级,模型性能显著提高,达新基准最优。
  • 适合想提升模型通用推理能力的研究者使用。

大语言模型(LLMs)在复杂推理方面取得了突破性进展。然而,现有强化学习研究多聚焦于特定领域任务(如数学或代码生成),难以覆盖真实世界中多样且复杂的推理场景。为此,我们提出InternBootcamp,一个开源框架,包含1000多个跨领域任务环境,专为大模型推理研究设计。基于此,我们进一步构建了自动化的评估基准Bootcamp-Eval,用于全面评估模型表现。评估发现,前沿模型在多数任务上仍表现不足。通过InternBootcamp训练,模型性能显著提升,我们的32B参数模型在Bootcamp-Eval上达到当前最佳,并在其他主流基准上同样表现出色。关键验证表明,性能提升主要来自任务规模的扩大——超过两个数量级的任务扩展,为打造具备通用推理能力的智能体提供了可行路径。所有数据与代码均已公开。

原文摘要 · Abstract (English)

Large language models (LLMs) have revolutionized artificial intelligence by enabling complex reasoning capabilities. While recent advancements in reinforcement learning (RL) have primarily focused on domain-specific reasoning tasks (e.g., mathematics or code generation), real-world reasoning scenarios often require models to handle diverse and complex environments that narrow-domain benchmarks cannot fully capture. To address this gap, we present InternBootcamp, an open-source framework comprising 1000+ domain-diverse task environments specifically designed for LLM reasoning research. With these bootcamps, we further establish Bootcamp-Eval, an automatically generated benchmark for comprehensive performance assessment. Evaluation reveals that frontier models still underperform in many reasoning tasks, while training with InternBootcamp provides an effective way to significantly improve performance, leading to our 32B model that achieves stateof-the-art results on Bootcamp-Eval and excels on other established benchmarks. In particular, we validate that consistent performance gain come from including more training tasks, namely task scaling, over two orders of magnitude, offering a promising route towards capable reasoning generalist. All data and code are publicly available.

大模型推理任务扩展开源框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。