提出内在深度思考技能,让大模型自己学会复杂推理。
HeavySkill: Heavy Thinking as the Inner Skill in Agentic Harness

- 将深度思考拆解为并行推理+摘要的两阶段机制
- 性能超越传统Best-of-N,强模型接近Pass@N效果
- 可通过强化学习扩展思考深度与广度,适合自进化模型
近期基于编排框架的智能体协同系统在复杂推理任务中取得显著进展,但其核心驱动力仍被复杂的系统设计所掩盖。本文提出HeavySkill视角,将深度思考视为编排框架中的最小执行单元,并作为内化于模型参数中的内在技能,驱动智能体完成复杂任务。该技能表现为并行推理后汇总的两阶段流程,可嵌入任意智能体框架。我们在多个领域开展系统性实证研究,结果表明该内在技能持续优于传统Best-of-N(BoN)策略;值得注意的是,更强的大模型甚至可逼近Pass@N表现。关键在于,通过强化学习可进一步扩展深度思考的深度与宽度,为无需依赖脆弱编排层的自进化大模型提供了可行路径。
原文摘要 · Abstract (English)
Recent advances in agentic harness with orchestration frameworks that coordinate multiple agents with memory, skills, and tool use have achieved remarkable success in complex reasoning tasks. However, the underlying mechanism that truly drives performance remains obscured behind intricate system designs. In this paper, we propose HeavySkill, a perspective that views heavy thinking not only as a minimal execution unit in orchestration harness but also as an inner skill internalized within the model's parameters that drives the orchestrator to solve complex tasks. We identify this skill as a two-stage pipeline, i.e., parallel reasoning then summarization, which can operate beneath any agentic harness. We present a systematic empirical study of HeavySkill across diverse domains. Our results show that this inner skill consistently outperforms traditional Best-of-N (BoN) strategies; notably, stronger LLMs can even approach Pass@N performance. Crucially, we demonstrate that the depth and width of heavy thinking, as a learnable skill, can be further scaled via reinforcement learning, offering a promising path toward self-evolving LLMs that internalize complex reasoning without relying on brittle orchestration layers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。