提出新基准与优化框架,公平评估命令行智能体的性能与效率。
Matching Matters: A Fair Quality-Efficiency Benchmark for Command-Line Agents
- 引入联合衡量质量与效率的AMS评分指标。
- 实验证明无通用最优命令行接口,不同组合表现差异显著。
- 通过描述优化实现更公平的跨模型评测,适合部署前评估。
大型语言模型的快速发展提升了基于命令行界面(CLI)智能体的任务求解能力,而CLI决定了模型调用工具、维护交互历史及故障恢复的方式。因此,有效匹配语言模型与CLI变得至关重要。然而,现有智能体评测多关注成功率,忽视成本、效率以及模型-指令接口组合选择等实际部署关键因素。为此,本文提出AgentMeter基准,引入新的代理评分(AMS),综合刻画任务质量、预算敏感性及资源密集型零奖励执行,实现对已部署模型-接口对的全面评估。此外,任务描述可能无意中偏好特定的模型-接口组合,导致评价结果反映描述特异性优势而非通用任务解决能力。为此,我们提出AgentMeter-Opt——一种基于轨迹的优化框架,构建适配组合且保持任务不变的描述变体,以建立更公平的评测集。大量实验表明:无单一接口在所有模型上均最优;任务成功率、执行成本与AMS识别出不同的最优配置。AgentMeter-Opt结果显示,任务保持的描述变化对不同模型-接口对影响不均,可改变其相对排序。二者共同为命令行智能体提供实用、公正且贴近部署需求的评估基础。
原文摘要 · Abstract (English)
Rapid advances in large language models have improved the task-solving capabilities of command-line-interface (CLI)-based agents, whose CLIs determine how models invoke tools, maintain interaction history, and recover from failures. Consequently, effective matching between CLIs and LLMs has become essential. However, existing agent benchmarks largely emphasize success rate while overlooking practical objectives such as cost and efficiency, as well as the selection of LM-CLI combinations, all of which are critical in real-world deployment. We therefore introduce AgentMeter, a quality-efficiency benchmark with a new metric, the AgentMeter Score (AMS), that jointly characterizes task quality, budget sensitivity, and resource-intensive zero-reward execution, enabling a more complete assessment of deployed LM-CLI pairs. Furthermore, collected task descriptions may inadvertently favor LM-CLI pairs that are particularly compatible with their wording and structure, causing evaluation results to reflect description-specific advantages rather than general task-solving capability. We therefore propose AgentMeter-Opt, a trajectory-grounded optimization framework that constructs pair-adapted, task-preserving description variants to build a fairer evaluation set across LM-CLI pairs. Extensive experiments show that no CLI is universally optimal across language models and that task success, execution cost, and AMS identify different competitive configurations. Results on AgentMeter-Opt further reveal that task-preserving description changes affect LM-CLI pairs unevenly and can alter their relative ordering across valid description conditions. Together, AgentMeter and AgentMeter-Opt provide a practical foundation for fair and deployment-relevant evaluation of command-line agents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。