构建首个覆盖多种孟加拉方言的自发语音语料库,助力方言自适应语音识别。
RegSpeech12: A Regional Corpus of Bengali Spontaneous Speech Across Dialects
- 采集并标注5大主方言及孟加拉国多个地区的自发语音数据
- 验证方言间音系与形态差异对自动语音识别的影响
- 适合方言研究、语音识别与数字包容性技术开发者
孟加拉语在南亚及海外侨民中广泛使用,呈现出由地理、文化与历史塑造的显著方言多样性。语音与发音特征可大致划分为五大主要方言群:东孟加拉语、曼布赫米语、朗普里语、瓦伦德拉语和拉尔希语。在孟加拉国境内,查塔贡、锡尔赫特、兰格普尔、拉杰沙希、诺阿哈利和巴里沙尔等地区还存在词汇、句法与形态上的进一步差异。尽管语言资源丰富,但针对孟加拉方言的计算处理系统性研究仍较匮乏。本研究旨在记录并分析这些方言的语音与形态特征,探索为区域性变体定制自动语音识别(ASR)系统的可行性。此类工作有望应用于虚拟助手等场景,促进方言多样性保护与包容性语言技术的发展。本研究构建的数据集已公开发布,供学术界使用。
原文摘要 · Abstract (English)
The Bengali language, spoken extensively across South Asia and among diasporic communities, exhibits considerable dialectal diversity shaped by geography, culture, and history. Phonological and pronunciation-based classifications broadly identify five principal dialect groups: Eastern Bengali, Manbhumi, Rangpuri, Varendri, and Rarhi. Within Bangladesh, further distinctions emerge through variation in vocabulary, syntax, and morphology, as observed in regions such as Chittagong, Sylhet, Rangpur, Rajshahi, Noakhali, and Barishal. Despite this linguistic richness, systematic research on the computational processing of Bengali dialects remains limited. This study seeks to document and analyze the phonetic and morphological properties of these dialects while exploring the feasibility of building computational models particularly Automatic Speech Recognition (ASR) systems tailored to regional varieties. Such efforts hold potential for applications in virtual assistants and broader language technologies, contributing to both the preservation of dialectal diversity and the advancement of inclusive digital tools for Bengali-speaking communities. The dataset created for this study is released for public use.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。