首个面向台湾普通话的语音大模型,支持实时对话交互。
Building a Taiwanese Mandarin Spoken Language Model: A First Attempt
- 采用解码器架构,实现端到端语音对话生成。
- 构建合成对话数据集,优化实时多轮交互性能。
- 适合研究方言语音模型与实时对话系统者参考。
本技术报告介绍了我们首次为台湾普通话构建语音大语言模型(Spoken LLM)的尝试,旨在实现多轮对话中的实时语音到语音交互。所提出的端到端模型采用仅解码器的Transformer架构,力求在保持对话流畅性的同时支持全双工能力,即允许同时说话与听觉接收。论文详细描述了训练过程,包括使用合成对话进行数据准备,并针对实时交互进行了调整。我们还开发了一个平台,用于评估多轮对话中对话连贯性与响应自然度。希望本报告的发布能为未来台湾普通话语音大模型的发展提供参考。
原文摘要 · Abstract (English)
This technical report presents our initial attempt to build a spoken large language model (LLM) for Taiwanese Mandarin, specifically tailored to enable real-time, speech-to-speech interaction in multi-turn conversations. Our end-to-end model incorporates a decoder-only transformer architecture and aims to achieve seamless interaction while preserving the conversational flow, including full-duplex capabilities allowing simultaneous speaking and listening. The paper also details the training process, including data preparation with synthesized dialogues and adjustments for real-time interaction. We also developed a platform to evaluate conversational fluency and response coherence in multi-turn dialogues. We hope the release of the report can contribute to the future development of spoken LLMs in Taiwanese Mandarin.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。