首个开放的泰语依善方言对话语音数据集,助力小语种语音技术发展。
Developing an Open Conversational Speech Corpus for the Isan Language
- 基于自然对话构建,包含口语、停顿、语码转换等真实语言现象。
- 解决依善语无标准拼写难题,制定兼顾准确与可计算性的转写规范。
- 适合小语种语音研究、包容性AI开发及方言语音建模方向学者使用。
本文介绍了首个面向泰国最广泛使用的地区方言——依善语的开放对话语音数据集的构建工作。与以往以朗读或脚本为主的语音语料库不同,该数据集由自然对话构成,真实呈现了口语表达、即兴语调、不流畅现象以及频繁的依善语与标准泰语之间的语码转换。构建此资源的一大挑战在于依善语缺乏标准化拼写体系,其词汇声调与泰语差异显著,导致书写方式差异大,影响转写指南设计,引发一致性、可用性与语言真实性问题。为此,我们制定了兼顾表征准确性与计算处理需求的实用转写协议。通过将该数据集作为开源资源发布,旨在推动包容性人工智能发展,支持低资源语言研究,并为对话语音建模中的语言与技术挑战提供基础支撑。
原文摘要 · Abstract (English)
This paper introduces the development of the first open conversational speech dataset for the Isan language, the most widely spoken regional dialect in Thailand. Unlike existing speech corpora that are primarily based on read or scripted speech, this dataset consists of natural speech, thereby capturing authentic linguistic phenomena such as colloquials, spontaneous prosody, disfluencies, and frequent code-switching with central Thai. A key challenge in building this resource lies in the lack of a standardized orthography for Isan. Current writing practices vary considerably, due to the different lexical tones between Thai and Isan. This variability complicates the design of transcription guidelines and poses questions regarding consistency, usability, and linguistic authenticity. To address these issues, we establish practical transcription protocols that balance the need for representational accuracy with the requirements of computational processing. By releasing this dataset as an open resource, we aim to contribute to inclusive AI development, support research on underrepresented languages, and provide a basis for addressing the linguistic and technical challenges inherent in modeling conversational speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。