通过中间层技能监督提升工具型大模型的推理可靠性
SCRIBE: Structured Mid-Level Supervision for Tool-Using Language Models
- 引入技能原型库实现结构化奖励建模,减少评估噪声
- 使Qwen3-4B在AIME25上准确率从43.3%提升至63.3%
- 适合研究自主工具使用与多步推理的开发者
训练可靠的工具增强型智能体仍面临重大挑战,主要源于多步推理中的信用分配难题。尽管过程级奖励模型前景广阔,但现有基于LLM的评判者常因缺乏细粒度、任务特定的评分标准,难以区分高层规划与低层执行。本文提出SCRIBE(Skill-Conditioned Reward with Intermediate Behavioral Evaluation),一种在新颖中间层抽象上干预的强化学习框架。SCRIBE将奖励建模建立在精心构建的技能原型库之上,将开放式的LLM评估转化为受约束的验证问题。通过将每个子目标路由至对应原型,奖励模型获得精确的结构化评分标准,显著降低奖励方差。实验表明,SCRIBE在多个推理与工具使用基准上达到最先进性能。尤其在AIME25上,使Qwen3-4B模型准确率从43.3%提升至63.3%,并大幅提高复杂多轮工具交互的成功率。进一步分析显示,抽象层级间存在协同演化,中层技能掌握始终先于有效高层规划行为的出现。最后,证明SCRIBE可与底层工具优化叠加,为更自主、可靠的工具使用智能体提供可扩展且互补的路径。
原文摘要 · Abstract (English)
Training reliable tool-augmented agents remains a significant challenge, largely due to the difficulty of credit assignment in multi-step reasoning. While process-level reward models offer a promising direction, existing LLM-based judges often produce noisy and inconsistent signals because they lack fine-grained, task-specific rubrics to distinguish high-level planning from low-level execution. In this work, we introduce SCRIBE (Skill-Conditioned Reward with Intermediate Behavioral Evaluation), a reinforcement learning framework that intervenes at a novel mid-level abstraction. SCRIBE grounds reward modeling in a curated library of skill prototypes, transforming open-ended LLM evaluation into a constrained verification problem. By routing each subgoal to a corresponding prototype, the reward model is equipped with precise, structured rubrics that substantially reduce reward variance. Experimental results show that SCRIBE achieves state-of-the-art performance across a range of reasoning and tool-use benchmarks. In particular, it improves the AIME25 accuracy of a Qwen3-4B model from 43.3% to 63.3%, and significantly increases success rates in complex multi-turn tool interactions. Further analysis of training dynamics reveals a co-evolution across abstraction levels, where mastery of mid-level skills consistently precedes the emergence of effective high-level planning behaviors. Finally, we demonstrate that SCRIBE is additive to low-level tool optimizations, providing a scalable and complementary pathway toward more autonomous and reliable tool-using agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。