通过输入重构提升大模型在动态环境中的工具使用准确率
How Can Input Reformulation Improve Tool Usage Accuracy in a Complex Dynamic Environment? A Study on $τ$-bench
- 将用户查询与领域规则、工具建议结合重构输入
- 在τ-bench上使通过率提升16.1%至19.1%
- 适合需要高可靠性的复杂对话系统开发
大型语言模型在推理和规划能力上的进展使其具备在动态环境中自主使用工具的潜力。然而,在多轮对话环境如τ-bench中,这些智能体常因持续推理不一致、违反领域规则及长期工具调用中信息提取错误而表现不佳。我们对对话轨迹中的常见错误进行了全面的手动分析,并实验了对工具调用智能体输入的重构策略。最终提出输入重构多智能体(IRMA)框架,自动将用户查询与相关领域规则和工具建议融合以优化输入。结果显示,IRMA在整体pass^5得分上分别优于ReAct、Function Calling和Self-Reflection 16.1%、12.7%和19.1%,表明其在动态环境下的可靠性与一致性显著更优。
原文摘要 · Abstract (English)
Recent advances in reasoning and planning capabilities of large language models (LLMs) have enabled their potential as autonomous agents capable of tool use in dynamic environments. However, in multi-turn conversational environments like $τ$-bench, these agents often struggle with consistent reasoning, adherence to domain-specific policies, and extracting correct information over a long horizon of tool-calls and conversation. To capture and mitigate these failures, we conduct a comprehensive manual analysis of the common errors occurring in the conversation trajectories. We then experiment with reformulations of inputs to the tool-calling agent for improvement in agent decision making. Finally, we propose the Input-Reformulation Multi-Agent (IRMA) framework, which automatically reformulates user queries augmented with relevant domain rules and tool suggestions for the tool-calling agent to focus on. The results show that IRMA significantly outperforms ReAct, Function Calling, and Self-Reflection by 16.1%, 12.7%, and 19.1%, respectively, in overall pass^5 scores. These findings highlight the superior reliability and consistency of IRMA compared to other methods in dynamic environments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。