arXiv:2604.27204cs.CLcs.LG2026-04中稿 · LREC 2026

通过语言间差异迁移,提升语音转写模型的准确性和新特征识别能力。

Selective Augmentation: Improving Universal Automatic Phonetic Transcription via G2P Bootstrapping

  • 选择性地从辅助语言(印地语)迁移发音差异,增强训练数据。
  • 辅音送气识别率从0%提升至61.2%,清浊音误报率显著下降。
  • 适合需要跨语言语音建模与特征扩展的研究者使用。

在通用自动音素转写(APT)领域,高质量且多样化的训练转写数据至关重要,但现有数据有限。本文提出一种名为选择性增强(Selective Augmentation)的自举方法,通过有选择地转移不同语言间的发音差异来提升训练数据质量。基于MultIPA模型,我们以印地语为辅助语言,成功在现有特征(清浊音)上将准确率提升17.6%(通过减少假阳性),并引入新特征(送气)。原本基线模型对德语 /p, t, k/ 的送气识别率为0%,而本方法达到61.2%。该改进使清音类别的占比降低32.2%,有效缓解了测试语言中塞音之间的混淆问题。

原文摘要 · Abstract (English)

In the field of universal automatic phonetic transcription (APT), clean and diverse training transcriptions are required. However, such high-quality data is limited. We propose the bootstrapping approach Selective Augmentation to improve the available training transcriptions by selectively transferring distinctions between languages. Based on the model MultIPA, we exemplarily show that we could increase the accuracy of an existing feature (plosive voicing) and add a new feature (plosive aspiration) by augmenting the existing training data using information from a separate helper language (Hindi). We describe intrinsic challenges of the evaluation and develop objective metrics to determine the success: Voicing accuracy was increased by 17.6% by reducing the number of false positives. Additionally, aspiration recognition was introduced: While the baseline transcribed 0% of German /p, t, k/ as aspirated, our approach transcribed them as aspirated in 61.2% of the cases. Introducing aspiration recognition to APT models allowed for the tenuis class to be successfully reduced by 32.2%, which also reduces the conflations between the test language's plosives.

语音转写跨语言数据增强音系特征

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。