构建德语方言ASR与翻译数据集,助力模型理解地方口音。
A Multi-Dialectal Dataset for German Dialect ASR and Dialect-to-Standard Speech Translation
- 收集东南德三类方言及标准德语共4.5小时读诵语音
- 方言转写与标准语转写并列,支持跨语言对比分析
- 发现模型更倾向保留方言结构而非强行标准化
尽管德国方言多样,但当前自动语音识别(ASR)研究仍严重缺乏方言数据。为评估模型对方言变异的鲁棒性,我们发布Betthupferl数据集,包含东南德三类方言(弗兰克尼亚、巴伐利亚、阿勒曼尼)共四小时读诵语音,以及半小时标准德语语音。所有音频均配有方言和标准德语双语转写,并分析其语言差异。我们在方言转标准语任务上测试多个先进多语言ASR模型,发现模型输出在接近方言或标准语转写之间存在差异。对最优模型的定性分析显示,其常保留方言语法结构,仅部分进行规范化处理。
原文摘要 · Abstract (English)
Although Germany has a diverse landscape of dialects, they are underrepresented in current automatic speech recognition (ASR) research. To enable studies of how robust models are towards dialectal variation, we present Betthupferl, an evaluation dataset containing four hours of read speech in three dialect groups spoken in Southeast Germany (Franconian, Bavarian, Alemannic), and half an hour of Standard German speech. We provide both dialectal and Standard German transcriptions, and analyze the linguistic differences between them. We benchmark several multilingual state-of-the-art ASR models on speech translation into Standard German, and find differences between how much the output resembles the dialectal vs. standardized transcriptions. Qualitative error analyses of the best ASR model reveal that it sometimes normalizes grammatical differences, but often stays closer to the dialectal constructions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。