arXiv:2410.17196cs.CLcs.AI2024-10Transactions of th…被引 215

首个评估大模型语音助手真实场景表现的基准测试

VoiceBench: Benchmarking LLM-Based Voice Assistants

  • 设计包含真实与合成语音指令的多维度评测框架
  • 发现现有语音助手在复杂环境下表现显著下降
  • 适合研究语音交互、大模型应用的开发者与学者

基于大语言模型(LLMs)的语音助手,如 GPT-4o,已实现实时语音交互,显著优于传统文本交互。然而,缺乏专门评估语音交互能力的基准测试,阻碍了该领域发展。现有评估多集中于清晰语音下的自动语音识别(ASR)或通用知识,忽略了真实世界中多样化的说话人特征、环境因素和内容复杂性。为此,我们提出 VoiceBench,首个针对 LLM 语音助手的多维度评估基准。该基准包含真实与合成语音指令,涵盖上述三类现实场景变量。大量实验揭示了当前语音助手模型的局限性,并为未来研究提供了重要启示。

原文摘要 · Abstract (English)

Building on the success of large language models (LLMs), recent advancements such as GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering a significantly improved user experience compared to traditional text-based interactions. However, the absence of benchmarks designed to evaluate these speech interaction capabilities has hindered progress of LLM-based voice assistants development. Current evaluations focus primarily on automatic speech recognition (ASR) or general knowledge evaluation with clean speeches, neglecting the more intricate, real-world scenarios that involve diverse speaker characteristics, environmental and content factors. To address this, we introduce VoiceBench, the first benchmark designed to provide a multi-faceted evaluation of LLM-based voice assistants. VoiceBench also includes both real and synthetic spoken instructions that incorporate the above three key real-world variations. Extensive experiments reveal the limitations of current LLM-based voice assistant models and offer valuable insights for future research and development in this field.

语音助手大模型评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。