arXiv:2602.22522cs.CLcs.AI2026-02中稿 · LREC 2026被引 2

提出新模型,让低资源客语语音识别更准,同时兼顾汉字和拼音。

Efficient Dialect-Aware Modeling and Conditioning for Low-Resource Taiwanese Hakka Speech Processing

  • 用RNN-T框架分离方言风格与语言内容,提升泛化能力
  • 在客家语数据集上,汉字和拼音识别错误率分别降低57%和40.41%
  • 首个统一处理客语方言与双文字系统的单模型,适合濒危语言研究

台湾客家话是一种低资源、濒危语言,其自动语音识别(ASR)面临显著挑战,包括高度的方言变异以及存在汉字和拼音两种书写系统。传统ASR模型在此背景下易将语言内容与方言特征混淆,影响识别效果。为此,本文提出基于循环神经网络转换器(RNN-T)的统一框架,引入方言感知建模策略,旨在解耦方言‘风格’与语言‘内容’,增强模型学习鲁棒且泛化的表示能力。此外,框架采用参数高效预测网络,同时建模汉字和拼音的ASR任务。实验表明,两项任务间形成强大协同效应,跨文字目标作为相互正则化项,有效提升主任务性能。在HAT语料库上的测试显示,该模型在汉字和拼音ASR上分别实现57.00%和40.41%的相对误差率降低。据我们所知,这是首次系统研究客家方言变异对ASR的影响,也是首个能联合处理此类任务的单一模型。

原文摘要 · Abstract (English)

Taiwanese Hakka is a low-resource, endangered language that poses significant challenges for automatic speech recognition (ASR), including high dialectal variability and the presence of two distinct writing systems (Hanzi and Pinyin). Traditional ASR models often encounter difficulties in this context, as they tend to conflate essential linguistic content with dialect-specific variations across both phonological and lexical dimensions. To address these challenges, we propose a unified framework grounded in the Recurrent Neural Network Transducers (RNN-T). Central to our approach is the introduction of dialect-aware modeling strategies designed to disentangle dialectal "style" from linguistic "content", which enhances the model's capacity to learn robust and generalized representations. Additionally, the framework employs parameter-efficient prediction networks to concurrently model ASR (Hanzi and Pinyin). We demonstrate that these tasks create a powerful synergy, wherein the cross-script objective serves as a mutual regularizer to improve the primary ASR tasks. Experiments conducted on the HAT corpus reveal that our model achieves 57.00% and 40.41% relative error rate reduction on Hanzi and Pinyin ASR, respectively. To our knowledge, this is the first systematic investigation into the impact of Hakka dialectal variations on ASR and the first single model capable of jointly addressing these tasks.

语音识别方言建模低资源语言双文字系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。