统一对话系统演示工具,一键对比语音交互性能
ESPnet-SDS: Unified Toolkit and Demo for Spoken Dialogue Systems
- 构建统一网页界面,支持多种端到端与分步式对话系统
- 实时评估延迟、理解力、回复质量与音频清晰度等指标
- 适合语音系统研究者快速比较技术优劣
音频基础模型的进步推动了端到端(E2E)语音对话系统的发展,但各系统独立的网页接口使得有效对比变得困难。为此,我们推出一个开源、易用的工具包,用于为各类级联式和端到端语音对话系统构建统一网页界面。演示系统还提供实时自动化评估指标,包括:(1)延迟,(2)对用户输入的理解能力,(3)系统回复的连贯性、多样性与相关性,(4)输出语音的可懂度与音频质量。基于人类-人类对话数据集作为代理,我们利用这些指标对比多种级联与端到端系统。分析表明,该工具包使研究人员能轻松对比不同技术,揭示当前端到端系统存在音频质量较差、回复多样性不足等问题。示例演示已公开:https://huggingface.co/spaces/Siddhant/Voice_Assistant_Demo。
原文摘要 · Abstract (English)
Advancements in audio foundation models (FMs) have fueled interest in end-to-end (E2E) spoken dialogue systems, but different web interfaces for each system makes it challenging to compare and contrast them effectively. Motivated by this, we introduce an open-source, user-friendly toolkit designed to build unified web interfaces for various cascaded and E2E spoken dialogue systems. Our demo further provides users with the option to get on-the-fly automated evaluation metrics such as (1) latency, (2) ability to understand user input, (3) coherence, diversity, and relevance of system response, and (4) intelligibility and audio quality of system output. Using the evaluation metrics, we compare various cascaded and E2E spoken dialogue systems with a human-human conversation dataset as a proxy. Our analysis demonstrates that the toolkit allows researchers to effortlessly compare and contrast different technologies, providing valuable insights such as current E2E systems having poorer audio quality and less diverse responses. An example demo produced using our toolkit is publicly available here: https://huggingface.co/spaces/Siddhant/Voice_Assistant_Demo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。