新基准评估对话智能体应对目标突变的适应能力,发现高准确率不等于真抗变。
AgentChangeBench: A Multi-Dimensional Evaluation Framework for Goal-Shift Robustness in Conversational AI
- 设计多维度评测框架,量化目标变更后任务成功率、工具使用效率等四指标。
- 测试显示GPT-4o在航班预订目标切换中恢复率达92.2%,而Gemini仅48.6%。
- 揭示模型冗余调用超80%且恢复延迟显著,适合企业级对话系统开发者参考。
目标变化是真实多轮交互的核心特征,但现有代理评测主要针对静态目标或一次性工具使用。我们提出AgentChangeBench,一个专为评估工具增强语言模型代理在三个企业领域中应对对话进行时目标变更的鲁棒性而设计的基准。该框架通过四个互补指标进行形式化评估:任务成功率(TSR)衡量有效性,工具使用效率(TUE)衡量可靠性,工具调用冗余率(TCRR)衡量浪费努力,目标变更恢复时间(GSRT)衡量适应延迟。AgentChangeBench包含2,835个任务序列和五个用户角色,每个角色均设计用于触发实际工作流中的转变点。在此设置下,我们评估了多个前沿模型,发现传统pass@k分数掩盖了显著差异:例如,GPT-4o在航班预订目标变更中达到92.2%恢复率,而Gemini下降至48.6%;零售任务虽参数有效性接近完美,但冗余率超过80%,暴露严重效率问题。这些发现表明,高原始准确率并不意味着动态目标下的鲁棒性,显式测量恢复时间和冗余度至关重要。AgentChangeBench为诊断和提升代理在真实企业场景中的韧性提供了可复现的测试平台。
原文摘要 · Abstract (English)
Goal changes are a defining feature of real world multi-turn interactions, yet current agent benchmarks primarily evaluate static objectives or one-shot tool use. We introduce AgentChangeBench, a benchmark explicitly designed to measure how tool augmented language model agents adapt to mid dialogue goal shifts across three enterprise domains. Our framework formalizes evaluation through four complementary metrics: Task Success Rate (TSR) for effectiveness, Tool Use Efficiency (TUE) for reliability, Tool Call Redundancy Rate (TCRR) for wasted effort, and Goal-Shift Recovery Time (GSRT) for adaptation latency. AgentChangeBench comprises 2,835 task sequences and five user personas, each designed to trigger realistic shift points in ongoing workflows. Using this setup, we evaluate several frontier models and uncover sharp contrasts obscured by traditional $\text{pass}@k$ scores: for example, GPT-4o reaches $92.2\%$ recovery on airline booking shifts while Gemini collapses to $48.6\%$, and retail tasks show near perfect parameter validity yet redundancy rates above $80\%$, revealing major inefficiencies. These findings demonstrate that high raw accuracy does not imply robustness under dynamic goals, and that explicit measurement of recovery time and redundancy is essential. AgentChangeBench establishes a reproducible testbed for diagnosing and improving agent resilience in realistic enterprise settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。