用丰富数据训练规则模型,让语音转换更快更准。
Fast, Not Fancy: Rethinking G2P with Rich Data and Rule-Based Models
- 用半自动流程构建专注同音字的数据集,提升标注效率
- 改进后的规则模型在波斯语上同音字识别准确率提升30%
- 适合对延迟敏感的无障碍应用,如屏幕阅读器
同音字消歧仍是音素转换(G2P)中的重大挑战,尤其在低资源语言中。该问题有两方面:一是构建平衡且全面的同音字数据集成本高、耗时长;二是特定消歧策略引入额外延迟,不适合屏幕阅读器等实时无障碍工具。本文提出解决方案:首先,设计半自动管道构建同音字聚焦数据集,推出HomoRich数据集,并用于提升最先进的深度学习型G2P系统在波斯语上的表现;其次,倡导范式转变——利用丰富的离线数据指导快速规则模型开发,适用于对延迟敏感的无障碍场景。为此,我们改进了最知名的规则型G2P系统eSpeak,构建出快速同音字感知版本HomoFast eSpeak。实验表明,深度学习模型与eSpeak系统的同音字消歧准确率均提升约30%。
原文摘要 · Abstract (English)
Homograph disambiguation remains a significant challenge in grapheme-to-phoneme (G2P) conversion, especially for low-resource languages. This challenge is twofold: (1) creating balanced and comprehensive homograph datasets is labor-intensive and costly, and (2) specific disambiguation strategies introduce additional latency, making them unsuitable for real-time applications such as screen readers and other accessibility tools. In this paper, we address both issues. First, we propose a semi-automated pipeline for constructing homograph-focused datasets, introduce the HomoRich dataset generated through this pipeline, and demonstrate its effectiveness by applying it to enhance a state-of-the-art deep learning-based G2P system for Persian. Second, we advocate for a paradigm shift - utilizing rich offline datasets to inform the development of fast, rule-based methods suitable for latency-sensitive accessibility applications like screen readers. To this end, we improve one of the most well-known rule-based G2P systems, eSpeak, into a fast homograph-aware version, HomoFast eSpeak. Our results show an approximate 30% improvement in homograph disambiguation accuracy for the deep learning-based and eSpeak systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。