arXiv:2509.26490cs.CLcs.AI2025-09被引 34

构建真实场景下多任务交互的智能体评测基准,挑战大模型实战能力。

VitaBench: Benchmarking LLM Agents with Versatile Interactive Tasks in Real-world Applications

  • 基于餐饮、购物、旅行等真实场景设计66种工具与100个跨场景任务
  • 顶尖模型在跨场景任务中成功率仅30%,单场景也低于50%
  • 支持动态交互与意图追踪,适合评估实用型智能体系统

随着基于大语言模型的智能体在真实场景中广泛应用,现有评测基准难以反映其处理海量信息、调用多样资源及应对动态用户交互的复杂性。为此,我们提出VitaBench,一个面向真实世界应用的复杂交互任务评测基准。该基准源自外卖、线下消费和在线旅游等日常服务场景,构建了迄今最复杂的生存类仿真环境,包含66种工具。通过去领域化策略框架,实现场景与工具的灵活组合,生成100个跨场景任务(主结果)和300个单场景任务。每项任务源于多个真实用户请求,要求智能体具备时空推理、复杂工具链使用、主动澄清模糊指令及多轮对话中跟踪用户意图的能力。此外,我们提出基于评分规则的滑动窗口评估器,可鲁棒评估复杂环境中多路径解决方案与随机交互表现。全面评估显示,即使最先进的模型在跨场景任务上成功率也仅为30%,其他任务低于50%。我们相信VitaBench将推动实际应用中智能体技术的发展。代码、数据集与排行榜已开源:https://vitabench.github.io/

原文摘要 · Abstract (English)

As LLM-based agents are increasingly deployed in real-life scenarios, existing benchmarks fail to capture their inherent complexity of handling extensive information, leveraging diverse resources, and managing dynamic user interactions. To address this gap, we introduce VitaBench, a challenging benchmark that evaluates agents on versatile interactive tasks grounded in real-world settings. Drawing from daily applications in food delivery, in-store consumption, and online travel services, VitaBench presents agents with the most complex life-serving simulation environment to date, comprising 66 tools. Through a framework that eliminates domain-specific policies, we enable flexible composition of these scenarios and tools, yielding 100 cross-scenario tasks (main results) and 300 single-scenario tasks. Each task is derived from multiple real user requests and requires agents to reason across temporal and spatial dimensions, utilize complex tool sets, proactively clarify ambiguous instructions, and track shifting user intent throughout multi-turn conversations. Moreover, we propose a rubric-based sliding window evaluator, enabling robust assessment of diverse solution pathways in complex environments and stochastic interactions. Our comprehensive evaluation reveals that even the most advanced models achieve only 30% success rate on cross-scenario tasks, and less than 50% success rate on others. Overall, we believe VitaBench will serve as a valuable resource for advancing the development of AI agents in practical real-world applications. The code, dataset, and leaderboard are available at https://vitabench.github.io/

智能体评测真实场景多轮交互大模型应用

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。