评测大模型在动态环境中的省钱规划与调整能力
CostBench: Evaluating Multi-Turn Cost-Optimal Planning and Adaptation in Dynamic Environments for LLM Tool-Use Agents
- 设计可扩展的成本导向基准,模拟真实世界变化
- 多任务下最优解匹配率不足75%,动态场景下降40%
- 适合关注经济决策与鲁棒性的智能体研究者
当前大模型智能体评估多聚焦任务完成度,忽视资源效率与适应性。为此,我们提出CostBench,一个面向旅行规划领域的可扩展、成本中心基准,用于评估智能体的经济推理与重规划能力。该基准包含多种原子与复合工具组合的任务,支持四类动态阻断事件(如工具失效、成本变动),以模拟现实不确定性并要求实时调整。在该基准上对主流开源及专有模型的评估显示:在静态环境下,智能体难以识别成本最优解,即使GPT-5在最困难任务上精确匹配率也低于75%;动态条件下性能进一步下降约40%。通过诊断这些缺陷,CostBench为构建更经济理性且稳健的下一代智能体奠定基础。
原文摘要 · Abstract (English)
Current evaluations of Large Language Model (LLM) agents primarily emphasize task completion, often overlooking resource efficiency and adaptability. This neglects a crucial capability: agents' ability to devise and adjust cost-optimal plans in response to changing environments. To bridge this gap, we introduce CostBench, a scalable, cost-centric benchmark designed to evaluate agents' economic reasoning and replanning abilities. Situated in the travel-planning domain, CostBench comprises tasks solvable via multiple sequences of atomic and composite tools with diverse, customizable costs. It also supports four types of dynamic blocking events, such as tool failures and cost changes, to simulate real-world unpredictability and necessitate agents to adapt in real time. Evaluating leading open-sourced and proprietary models on CostBench reveals a substantial gap in cost-aware planning: agents frequently fail to identify cost-optimal solutions in static settings, with even GPT-5 achieving less than 75% exact match rate on the hardest tasks, and performance further dropping by around 40% under dynamic conditions. By diagnosing these weaknesses, CostBench lays the groundwork for developing future agents that are both economically rational and robust.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。