arXiv:2605.13841cs.SDcs.AI2026-05被引 5

新基准EVA-Bench可自动评估语音助手的对话真实性和全链路表现。

EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents

论文配图:EVA-Bench: A New End-to-end Framework for Evaluating Voice Agents
图 1 · 摘自论文原文
  • 构建端到端仿真对话系统,自动验证并重生成对话以保证质量。
  • 提出双指标体系:任务完成度与对话体验,覆盖语音交互全场景。
  • 适用于各类语音助手架构,适合研发和评测人员使用。

语音助手是完成任务的智能对话系统,在企业应用中日益普及。然而,现有基准未能同时解决两大核心挑战:生成逼真模拟对话,以及全面衡量语音特有失败模式。我们提出EVA-Bench,一个端到端评估框架,兼顾模拟与测量。在模拟方面,通过动态多轮对话实现机器人间音频交互,并自动检测用户模拟器错误,及时重生成对话再评分。在测量方面,引入两个综合指标:EVA-A(准确性),涵盖任务完成、忠实度及语音保真度;EVA-X(体验),涵盖对话进展、表达简洁性与轮次时间。两者适用于所有主流代理架构,支持跨架构直接比较。EVA-Bench包含三个企业领域共213个场景,可控扰动集用于测试口音与噪声鲁棒性,以及pass@1、pass@k、pass^k等指标以区分峰值与可靠能力。对12种系统(涵盖全部三种架构)测试发现:(1) 无一系统在EVA-A pass@1与EVA-X pass@1上同时超过0.5;(2) 峰值与稳定性能差异显著,EVA-A中中位pass@k–pass^k差距达0.44;(3) 口音与噪声扰动暴露明显鲁棒性缺口,影响因架构、系统与指标而异,平均Δ高达0.314。完整框架、评测套件与数据已开源发布。

原文摘要 · Abstract (English)

Voice agents, artificial intelligence systems that conduct spoken conversations to complete tasks, are increasingly deployed across enterprise applications. However, no existing benchmark jointly addresses two core evaluation challenges: generating realistic simulated conversations, and measuring quality across the full scope of voice-specific failure modes. We present EVA-Bench, an end-to-end evaluation framework that addresses both. On the simulation side, EVA-Bench orchestrates bot-to-bot audio conversations over dynamic multi-turn dialogues, with automatic simulation validation that detects user simulator error and appropriately regenerates conversations before scoring. On the measurement side, EVA-Bench introduces two composite metrics: EVA-A (Accuracy), capturing task completion, faithfulness, and audio-level speech fidelity; and EVA-X (Experience), capturing conversation progression, spoken conciseness, and turn-taking timing. Both metrics apply to all major agent architectures, enabling direct cross-architecture comparison. EVA-Bench includes 213 scenarios across three enterprise domains, a controlled perturbation suite for accent and noise robustness, and pass@1, pass@k, pass^k measurements that distinguish peak from reliable capability. Across 12 systems spanning all three architectures, we find: (1) no system simultaneously exceeds 0.5 on both EVA-A pass@1 and EVA-X pass@1; (2) peak and reliable performance diverge substantially (median pass@k--pass^k gap of 0.44 on EVA-A); and (3) accent and noise perturbations expose substantial robustness gaps, with effects varying across architectures, systems, and metrics (mean $Δ$ up to 0.314). We release the full framework, evaluation suite, and benchmark data under an open-source license.

语音助手评估基准对话系统端到端

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。