用历史执行案例指导大模型工具使用,自动调节思考深度与结构正确性。
Case-Based Calibration of Adaptive Reasoning and Execution for LLM Tool Use

- 基于历史执行轨迹提取复杂度与失败信号,指导策略选择。
- 在BFCLv2和ToolBench上提升5.85%执行准确率,推理长度减少26%。
- 适合需要高可靠工具调用的自动化系统开发者参考。
工具使用拓展了大语言模型的能力边界,但可靠执行需平衡推理深度与结构有效性。本文从案例驱动视角提出CAST框架,将历史执行轨迹视为结构化案例。不同于直接复用原始输出,CAST提取案例衍生信号,用于识别任务复杂度以估计最优推理策略,并构建失败模式以预测结构失效点。该知识被转化为细粒度奖励设计与自适应推理机制,使模型在强化学习中自主内化案例策略。在BFCLv2和ToolBench上的实验表明,CAST不仅提升了方案符合性与任务成功率,还减少了冗余思考。相较基线,整体执行准确率最高提升5.85个百分点,平均推理长度降低26%,显著缓解了高影响结构错误。结果证明历史执行案例可作为可复用的校准知识,支持精准工具使用。
原文摘要 · Abstract (English)
Tool use extends large language models beyond parametric knowledge, but reliable execution requires balancing appropriate reasoning depth with strict structural validity. We approach this problem from a case-based perspective to present CAST, a case-driven framework that treats historical execution trajectories as structured cases. Instead of reusing raw exemplar outputs, CAST extracts case-derived signals to identify complexity profiles for estimating optimal reasoning strategies, alongside failure profiles to map likely structural breakdowns. The framework translates this knowledge into a fine-grained reward design and adaptive reasoning, enabling the model to autonomously internalize case-based strategies during reinforcement learning. Experiments on BFCLv2 and ToolBench demonstrate that CAST improves both schema-faithful execution and task-level tool-use success while reducing unnecessary deliberation. The approach achieves up to 5.85 percentage points gain in overall execution accuracy and reduces average reasoning length by 26%, significantly mitigating high-impact structural errors. Ultimately, this demonstrates how historical execution cases can provide reusable adaptation knowledge for calibrated tool use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。