构建中文方言与普通话语义对齐的语音表示,助力方言转普通话大模型发展。
Towards Comprehensive Semantic Speech Embeddings for Chinese Dialects
- 仅用语音识别数据训练语音编码器,实现方言与普通话的语义对齐。
- 在自建方言语音检索基准上达到领先性能,方言语音识别效果优异。
- 适合关注中文语音多语言建模、方言数字化的研究者和开发者。
尽管中文方言拥有数亿使用者,但在语音与语言技术方面仍远落后于普通话。多数方言以口语为主,因此构建方言到普通话的语音大语言模型(speech-LLM)比直接构建方言大语言模型更具实用性。实现这一目标的关键在于建立方言与普通话之间的跨方言语义对齐语音表示。本文通过仅使用语音识别(ASR)数据训练语音编码器,实现了这一对齐目标,并在我们新提出的汉语方言语音基准上验证了其在语音到语音检索任务中的有效性。该语音编码器在多个中文方言上的自动语音识别性能达到当前最优水平。我们的方言语音基准、语义对齐的语音表示以及语音到语音检索评估方法,为未来中文方言语音大模型的发展奠定了基础。相关数据集已开源:https://github.com/kalvinchang/yubao。
原文摘要 · Abstract (English)
Despite having hundreds of millions of speakers, Chinese dialects lag behind Mandarin in speech and language technologies. Most varieties are primarily spoken, making dialect-to-Mandarin speech-LLMs (large language models) more practical than dialect LLMs. Building dialect-to-Mandarin speech-LLMs requires speech representations with cross-dialect semantic alignment between Chinese dialects and Mandarin. In this paper, we achieve such a cross-dialect semantic alignment by training a speech encoder with ASR (automatic speech recognition)-only data, as demonstrated by speech-to-speech retrieval on a new benchmark of spoken Chinese varieties that we contribute. Our speech encoder further demonstrates state-of-the-art ASR performance on Chinese dialects. Together, our Chinese dialect benchmark, semantically aligned speech representations, and speech-to-speech retrieval evaluation lay the groundwork for future Chinese dialect speech-LLMs. We release the benchmark at https://github.com/kalvinchang/yubao.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。