arXiv:2607.14846cs.SDcs.AI2026-07被引 1

提出真实世界语音评测基准,揭示语音系统在多维度表现差异。

RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

  • 构建多维度评测体系,覆盖语音合成、转写、理解与识别
  • 实测显示各系统在自然度、情感表达等维度表现不一
  • 适合评估语音模型在真实场景下的综合能力

当前语音AI评测通常只关注语音可懂性、词错误率或文本对话质量等孤立能力,却很少检验系统是否利用了口语与文字表征之间的声学差异。为此,我们引入真实世界语音评测基准(RW-Voice-EQ Bench),用于评估文本到语音(TTS)、语音到语音(STS)、语音理解(SU)和自动语音识别(ASR)四类任务。评估结果表明性能高度依赖具体维度:对于TTS,自然度、表现力、身份稳定性和可靠性彼此独立;对于STS,拥有音频并不意味着使用语调情感,部分模型仍以文本驱动为主;对于SU,模型在副语言任务中表现参差;对于ASR,真实语境中的口音、情绪、噪声和对话条件暴露了传统干净语音评测未捕捉的失败。这些发现表明,语音AI应作为声学、表现力、交互性和鲁棒性能力的综合画像来评估,而非单一总分。

原文摘要 · Abstract (English)

Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation. To this end, we introduce the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech (TTS), speech-to-speech (STS), speech understanding (SU), and automatic speech recognition (ASR). Our evaluations indicate that performance is highly dimension-specific. For TTS, naturalness, expressiveness, identity stability, and reliability are largely independent evaluation dimensions. For STS, access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven. For SU, models perform unevenly across paralinguistic tasks. For ASR, real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks. Together, these results show that voice AI should be evaluated as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score.

语音评测语音生成真实场景多维度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。