arXiv:2412.01078cs.CLcs.AI2024-12被引 5

用6万小时合成语音数据训练中文语音大模型,实现低延迟实时对话。

Advancing Speech Language Models by Scaling Supervised Fine-Tuning with Over 60,000 Hours of Synthetic Speech Dialogue Data

  • 基于700万条中英双语合成对话数据,构建大规模语音语言模型。
  • 模型支持实时交互,延迟低、流畅度高,显著提升用户体验。
  • 专为中文场景设计,填补中文语音大模型研究空白,适合语音交互开发者。

GPT-4o在通过语音实现实时交互方面取得重要进展,其极低延迟与高流畅性不仅引人注目,也激发了该领域的研究兴趣。此类实时语音交互在需要快速反馈和即时响应的场景中尤为关键,能显著提升用户体验。然而,针对实时大语音语言模型的研究,特别是中文方向仍显不足。本文提出KE-Omni,一个基于大规模高质量合成语音对话数据集Ke-SpeechChat的无缝大语音语言模型,该数据集包含700万条中英文对话,涵盖42,002名说话者,总时长超过60,000小时,对推动该领域研究与发展具有重要意义。演示可访问:\url{https://huggingface.co/spaces/KE-Team/KE-Omni}。

原文摘要 · Abstract (English)

The GPT-4o represents a significant milestone in enabling real-time interaction with large language models (LLMs) through speech, its remarkable low latency and high fluency not only capture attention but also stimulate research interest in the field. This real-time speech interaction is particularly valuable in scenarios requiring rapid feedback and immediate responses, dramatically enhancing user experience. However, there is a notable lack of research focused on real-time large speech language models, particularly for Chinese. In this work, we present KE-Omni, a seamless large speech language model built upon Ke-SpeechChat, a large-scale high-quality synthetic speech interaction dataset consisting of 7 million Chinese and English conversations, featuring 42,002 speakers, and totaling over 60,000 hours, This contributes significantly to the advancement of research and development in this field. The demos can be accessed at \url{https://huggingface.co/spaces/KE-Team/KE-Omni}.

语音大模型实时交互合成数据中文语音

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。