arXiv:2603.28569cs.LGcs.AI2026-03被引 1

用真实云服务工单评估大模型代理,关注效率与真实表现。

CirrusBench: Evaluating LLM-based Agents Beyond Correctness in Real-World Cloud Service Environments

论文配图:CirrusBench: Evaluating LLM-based Agents Beyond Correctness in Real-World Cloud Service Environments
图 1 · 摘自论文原文
  • 基于真实云服务工单构建多轮交互评测框架
  • 发现顶尖模型在复杂任务中效率不足,难以满足客服要求
  • 适合研究智能客服、大模型应用落地的开发者

大语言模型(LLM)的代理能力日益增强,已应用于云服务等实际场景,客户交互具有高度技术复杂性和长程依赖性,鲁棒性和解决效率对客户满意度至关重要。然而,现有LLM代理评测基准多依赖合成环境,无法反映真实客户输入的多样性与不可预测性,且常忽略实际部署所需的关键效率指标。为此,我们提出CirrusBench,一个基于真实云服务工单数据的新型评测框架,保留技术场景中复杂的多轮逻辑链和真实工具依赖关系。超越执行正确性,引入以客户为中心的新指标,如归一化效率指数和多轮延迟,显式衡量解决效率。实验表明,尽管先进模型具备强推理能力,但在复杂真实的多轮任务中频繁表现不佳,未达到客服所需的高效率标准,揭示了未来实用技术客服应用的发展方向。CirrusBench评测框架已开源:https://github.com/CirrusAI

原文摘要 · Abstract (English)

The increasing agentic capabilities of Large Language Models (LLMs) have enabled their deployment in real-world applications, such as cloud services, where customer-assistant interactions exhibit high technical complexity and long-horizon dependencies, making robustness and resolution efficiency critical for customer satisfaction. However, existing benchmarks for LLM-based agents largely rely on synthetic environments that fail to capture the diversity and unpredictability of authentic customer inputs, often ignoring the resolution efficiency essential for real-world deployment. To bridge this gap, we introduce CirrusBench, a novel evaluation framework distinguished by its foundation in real-world data from authentic cloud service tickets. CirrusBench preserves the intricate multi-turn logical chains and realistic tool dependencies inherent to technical service environments. Moving beyond execution correctness, we introduce novel Customer-Centric metrics to define agent success, quantifying service quality through metrics such as the Normalized Efficiency Index and Multi-Turn Latency to explicitly measure resolution efficiency. Experiments utilizing our framework reveal that while state-of-the-art models demonstrate strong reasoning capabilities, they frequently struggle in complex, realistic multi-turn tasks and fail to meet the high-efficiency standards required for customer service, highlighting critical directions for the future development of LLM-based agents in practical technical service applications. CirrusBench evaluation framework is released at: https://github.com/CirrusAI

大模型代理评测基准智能客服效率评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。