开源工具包,让语音语言模型开发更简单高效
ESPnet-SpeechLM: An Open Speech Language Model Toolkit
- 将语音任务统一为序列建模,全流程标准化
- 支持1.7亿参数模型预训练,多任务表现优异
- 适合语音研究者与开发者快速构建语音智能应用
我们提出ESPnet-SpeechLM,一个开源工具包,旨在推动语音语言模型(SpeechLM)和语音驱动智能体应用的普及。该工具包通过将语音处理任务统一为通用的序列建模问题,提供从数据预处理、预训练、推理到任务评估的一体化工作流。用户可轻松定义任务模板并配置关键参数,实现流畅高效的SpeechLM开发。工具包在每个阶段均提供高度可配置模块,保障灵活性、效率与可扩展性。为展示其能力,我们提供了多个案例,包括在文本与语音任务上预训练的1.7B参数模型,在多种基准测试中表现优异。工具包及训练配方完全透明可复现,地址:https://github.com/espnet/espnet/tree/speechlm。
原文摘要 · Abstract (English)
We present ESPnet-SpeechLM, an open toolkit designed to democratize the development of speech language models (SpeechLMs) and voice-driven agentic applications. The toolkit standardizes speech processing tasks by framing them as universal sequential modeling problems, encompassing a cohesive workflow of data preprocessing, pre-training, inference, and task evaluation. With ESPnet-SpeechLM, users can easily define task templates and configure key settings, enabling seamless and streamlined SpeechLM development. The toolkit ensures flexibility, efficiency, and scalability by offering highly configurable modules for every stage of the workflow. To illustrate its capabilities, we provide multiple use cases demonstrating how competitive SpeechLMs can be constructed with ESPnet-SpeechLM, including a 1.7B-parameter model pre-trained on both text and speech tasks, across diverse benchmarks. The toolkit and its recipes are fully transparent and reproducible at: https://github.com/espnet/espnet/tree/speechlm.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。