arXiv:2510.21244cs.AI2025-10被引 3

构建专业外呼场景评估基准,提升大模型真实业务表现测试能力

VoiceAgentEval: A Dual-Dimensional Benchmark for Expert-Level Intelligent Voice-Agent Evaluation of Xbench's Professional-Aligned Series

  • 六领域30子场景分层设计,结合任务分解与领域适配评分
  • 大模型驱动的虚拟用户模拟真实情绪与沟通风格,增强测试真实性
  • 动态评估融合自动与人工判断,全面衡量专业能力与交互体验

我们提出OutboundEval,一个面向专家级智能外呼场景的大语言模型评估综合基准。针对现有方法在数据多样性、用户模拟真实性和评估指标准确性方面的三大缺陷,本研究构建了涵盖六大业务领域和30个代表性子场景的评测框架,每个场景包含特定流程拆解、加权评分和领域自适应指标。我们开发了基于大模型的用户模拟器,可生成具有多样化人格特征、真实行为模式、情感变化和沟通风格的虚拟用户,提供可控且真实的测试环境。同时引入动态评估机制,根据任务变化灵活调整,融合自动化评估与人工介入,全面衡量任务完成精度、专业知识应用、应变能力及用户体验质量。对12个先进大模型的实验揭示了专家级任务完成度与交互流畅性之间的显著权衡,为打造可靠、类人化外呼AI系统提供了实践指导。OutboundEval确立了可扩展、面向专业应用的标准化评估范式。

原文摘要 · Abstract (English)

We propose OutboundEval, a comprehensive benchmark for evaluating large language models (LLMs) in expert-level intelligent outbound calling scenarios. Unlike existing methods that suffer from three key limitations - insufficient dataset diversity and category coverage, unrealistic user simulation, and inaccurate evaluation metrics - OutboundEval addresses these issues through a structured framework. First, we design a benchmark spanning six major business domains and 30 representative sub-scenarios, each with scenario-specific process decomposition, weighted scoring, and domain-adaptive metrics. Second, we develop a large-model-driven User Simulator that generates diverse, persona-rich virtual users with realistic behaviors, emotional variability, and communication styles, providing a controlled yet authentic testing environment. Third, we introduce a dynamic evaluation method that adapts to task variations, integrating automated and human-in-the-loop assessment to measure task execution accuracy, professional knowledge application, adaptability, and user experience quality. Experiments on 12 state-of-the-art LLMs reveal distinct trade-offs between expert-level task completion and interaction fluency, offering practical insights for building reliable, human-like outbound AI systems. OutboundEval establishes a practical, extensible, and domain-oriented standard for benchmarking LLMs in professional applications.

大模型评估智能外呼虚拟用户动态测评

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。