对比文本与语音模型在德语方言间的迁移效果,发现语音模型更适合方言数据。
Standard-to-Dialect Transfer Trends Differ across Text and Speech: A Case Study on Intent and Topic Classification in German Dialects
- 比较文本、语音及语音转写后处理三种迁移方式
- 语音模型在方言意图分类上表现最佳,文本模型在标准语上最优
- 语音转写若输出规范文本,可显著提升方言处理效果
从标准语向非标准方言迁移的研究多集中于文本数据,但方言主要以口语形式存在,非标准拼写会影响文本处理。本文比较了三种设置下的标准语到方言迁移:纯文本模型、纯语音模型,以及先语音转写再由文本模型处理的级联系统。研究聚焦德语方言中的书面和口语意图分类,并发布首个德语方言语音意图分类数据集,同时开展话题分类实验。结果表明,语音模型在方言数据上表现最佳,而文本模型在标准语数据上更优;尽管级联系统整体落后于纯文本模型,但如果语音转写系统输出标准化的文本,其在方言数据上的表现仍较理想。
原文摘要 · Abstract (English)
Research on cross-dialectal transfer from a standard to a non-standard dialect variety has typically focused on text data. However, dialects are primarily spoken, and non-standard spellings cause issues in text processing. We compare standard-to-dialect transfer in three settings: text models, speech models, and cascaded systems where speech first gets automatically transcribed and then further processed by a text model. We focus on German dialects in the context of written and spoken intent classification -- releasing the first dialectal audio intent classification dataset -- with supporting experiments on topic classification. The speech-only setup provides the best results on the dialect data while the text-only setup works best on the standard data. While the cascaded systems lag behind the text-only models for German, they perform relatively well on the dialectal data if the transcription system generates normalized, standard-like output.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。