arXiv:2608.26546cs.AIcs.CL2026-08

真实工作流中评估智能体,发现大模型与框架共同决定其鲁棒性。

DuMateBench: Evaluating Autonomous Agents in Complex Real-World Workflows

  • 基于真实用户会话重建复杂任务,保留历史状态与环境扰动。
  • 200个任务跨8类场景,多数需多能力协同完成,性能差距显著。
  • 适合研究智能体系统性鲁棒性与框架-模型协同设计的学者。

自主智能体在现实世界中越来越多地用于完成复杂的多工具工作流。然而,现有基准通常按应用或能力分离任务,且评估环境比实际更清洁、更稳定。我们提出了 DuMateBench,一个从大规模生产级智能体平台匿名化并隐私保护后的用户会话中重建的真实会话基准。每个任务保留相关前置交互历史、持久配置和工作区状态,并经人工验证。该基准包含200个任务,覆盖8个广泛场景和17个细粒度能力类别,大多数任务需多能力协调。我们在隔离的Docker容器中注入三种真实环境复杂性(不足、不稳定、嘈杂),采用混合确定性与大模型作为裁判的评估协议进行测试。五种代表性智能体框架搭配四种先进大模型的实验揭示了严格任务完成度的巨大差距。补充的鲁棒性、效率与诊断分析表明,环境扰动下的表现由大模型能力和周围智能体框架共同决定。代码与数据公开于 https://dumatebench.com/。

原文摘要 · Abstract (English)

Autonomous agents are increasingly adopted to complete complex, multi-tool workflows in real-world settings. However, existing benchmarks typically separate tasks by application or capability and evaluate agents in environments that are cleaner and more stable than those encountered in practice. We introduce DuMateBench, a real-session benchmark reconstructed from anonymized and privacy-screened user sessions collected from a large-scale production agent platform. Each task preserves the relevant pre-solution interaction history, persistent configurations, and workspace state, and is then validated through human verification. The resulting benchmark comprises 200 tasks spanning 8 broad scenarios and 17 fine-grained capability categories, with most tasks requiring multiple capability coordination. We execute these tasks in isolated Docker containers injected with three forms of real-world environmental complexity: Insufficient, Unstable, and Noisy, and assess performance using a hybrid deterministic and LLM-as-Judge evaluation protocol. Experiments across five representative autonomous-agent frameworks paired with four state-of-the-art LLMs reveal substantial gaps in strict task completion. Complementary robustness, efficiency, and diagnostic analyses further show that performance under environmental perturbations is jointly shaped by the capabilities of the LLM and the surrounding agent framework. The code and data are publicly available at https://dumatebench.com/.

智能体评估真实场景工作流大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。