让大模型在企业中稳定执行多步骤任务,准确率从30%+提升至95%以上。
Routine: A Structural Planning Framework for LLM Agent System in Enterprise
- 设计结构化规划框架,明确指令与参数传递,提升任务执行稳定性。
- GPT-4o工具调用准确率从41.1%提升至96.3%,Qwen3-14B从32.6%升至83.3%。
- 适用于需要高可靠性的企业级智能流程自动化场景。
企业在部署代理系统时常面临挑战:通用模型缺乏领域特定流程知识,导致计划混乱、关键工具缺失、执行不稳定。本文提出Routine,一种多步规划框架,通过清晰的结构、明确的指令和无缝的参数传递,引导代理模块高效完成多步工具调用任务。在真实企业场景评估中,Routine使GPT-4o的工具调用准确率从41.1%提升至96.3%,Qwen3-14B从32.6%提升至83.3%。我们构建了遵循Routine的训练数据并微调Qwen3-14B,使其在特定场景下准确率达88.2%,体现更强的计划遵循能力。进一步基于Routine进行知识蒸馏,生成特定场景的多步工具调用数据集,微调后模型准确率达到95.5%,接近GPT-4o水平。结果表明Routine能有效提炼领域特定工具使用模式,增强模型对新场景的适应性。实验验证了Routine在构建稳定代理工作流方面的实用价值,加速企业级代理系统的落地与应用,推动AI for Process的技术愿景。
原文摘要 · Abstract (English)
The deployment of agent systems in an enterprise environment is often hindered by several challenges: common models lack domain-specific process knowledge, leading to disorganized plans, missing key tools, and poor execution stability. To address this, this paper introduces Routine, a multi-step agent planning framework designed with a clear structure, explicit instructions, and seamless parameter passing to guide the agent's execution module in performing multi-step tool-calling tasks with high stability. In evaluations conducted within a real-world enterprise scenario, Routine significantly increases the execution accuracy in model tool calls, increasing the performance of GPT-4o from 41.1% to 96.3%, and Qwen3-14B from 32.6% to 83.3%. We further constructed a Routine-following training dataset and fine-tuned Qwen3-14B, resulting in an accuracy increase to 88.2% on scenario-specific evaluations, indicating improved adherence to execution plans. In addition, we employed Routine-based distillation to create a scenario-specific, multi-step tool-calling dataset. Fine-tuning on this distilled dataset raised the model's accuracy to 95.5%, approaching GPT-4o's performance. These results highlight Routine's effectiveness in distilling domain-specific tool-usage patterns and enhancing model adaptability to new scenarios. Our experimental results demonstrate that Routine provides a practical and accessible approach to building stable agent workflows, accelerating the deployment and adoption of agent systems in enterprise environments, and advancing the technical vision of AI for Process.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。