用发音特征提升语音识别鲁棒性,实现零样本跨语言音素识别
ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition
- 基于发音特征构建预测框架,增强声学特征的通用性
- 在7种未见语言上使音素错误率降低20.56%,发音特征错误率降7.01%
- 适合做跨语言语音识别与低资源语言建模的研究者参考
零样本跨语言音素识别常因直接声学到符号映射的脆弱性而受阻,该映射易受语言特异性变化影响。受视觉领域联合嵌入预测架构(JEPA)启发,本文提出ArtNet,一种基于发音特征的结构化特征预测框架,以提升声学鲁棒性。ArtNet将发音预测器(从自监督学习特征中提取通用发音表征)与变分信息瓶颈(VIB)结合,抑制语言特异性差异。在7种未见语言上的实验表明,尤其当结合提出的向量空间库存对齐(VSIA)策略时,ArtNet显著优于现有基线,音素错误率(PER)相对降低20.56%,发音特征错误率(PFER)降低7.01%。
原文摘要 · Abstract (English)
Zero-shot cross-lingual phoneme recognition is often hindered by the fragility of direct acoustic-to-symbol mapping, which is susceptible to language-specific variations. Echoing joint-embedding predictive architecture (JEPA) work in vision, we propose ArtNet, a framework that explores a structured feature prediction task based on articulatory features to enhance acoustic robustness. Specifically, ArtNet integrates an articulatory predictor, designed to extract universal articulatory representations from self-supervised learning (SSL) features, with a variational information bottleneck (VIB) to suppress language-specific variations. Experiments on seven unseen languages demonstrate that ArtNet, particularly when synergized with the proposed vector-space inventory alignment (VSIA) strategy, significantly outperforms competitive baselines, achieving a 20.56\% relative reduction in phoneme error rate (PER) and 7.01\% in phoneme feature error rate (PFER).
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。