首个覆盖多语言、多轮对话与语音特质的语音对话模型评测基准
URO-Bench: Towards Comprehensive Evaluation for End-to-End Spoken Dialogue Models
- 构建双难度层级的S2S评测框架,覆盖理解、推理与口语能力
- 开源模型在日常问答表现尚可,指令遵循与语音信息理解仍不足
- 适合研究语音对话系统、多模态交互与语音认知的团队使用
大语言模型的进展推动了端到端语音对话模型(SDMs)的发展。与文本型大模型不同,语音对话模型的评估需兼顾认知维度(如逻辑推理、知识运用)和语音相关特性(如副语言线索、音频质量)。然而,目前在语音到语音(S2S)场景中仍缺乏全面的评测体系。为此,我们提出URO-Bench,首个涵盖多语言、多轮对话与副语言特性的S2S评测基准。该基准分为基础赛道与专业赛道,各含20个测试集,评估模型在理解、推理与口语交流方面的能力。评测结果显示,当前开源SDMs在日常问答任务中表现良好,但在指令遵循能力上逊于其底层大模型,并存在灾难性遗忘问题;在副语言信息与音频理解等高级评测中表现依然不佳,凸显该方向仍有待深入研究。URO-Bench旨在通过多维度评估促进语音对话模型的发展,助力跟踪领域进展。
原文摘要 · Abstract (English)
Recent advances in large language models (LLMs) have driven significant progress in end-to-end spoken dialogue models (SDMs). In contrast to text-based LLMs, the evaluation framework for SDMs should encompass both cognitive dimensions (e.g., logical reasoning, knowledge) and speech-related aspects (e.g., paralinguistic cues, audio quality). However, there is still a lack of comprehensive evaluations for SDMs in speech-to-speech (S2S) scenarios. To address this gap, we propose URO-Bench, an extensive benchmark for SDMs. Notably, URO-Bench is the first S2S benchmark that covers evaluations about multilingualism, multi-round dialogues, and paralinguistics. Our benchmark is divided into two difficulty levels: basic track and pro track, each comprising 20 test sets, evaluating the spoken dialogue model's abilities in Understanding, Reasoning, and Oral conversation. Evaluations on our proposed benchmark reveal that current open-source SDMs perform rather well in daily QA tasks, but lag behind their backbone LLMs in terms of instruction-following ability and also suffer from catastrophic forgetting. Their performance in advanced evaluations of paralinguistic information and audio understanding remains subpar, highlighting the need for further research in this direction. We hope that URO-Bench can facilitate the development of spoken dialogue models by providing a multifaceted evaluation of existing models and helping to track progress in this area.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。