arXiv:2607.23606cs.SDeess.AS2026-07中稿 · Interspeech 2026

用连续发音特征提升零样本语音分类,尤其改善罕见音素识别

Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features

论文配图:Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features
图 1 · 摘自论文原文
  • 用帧级发音特征向量替代离散音标进行分类
  • 对汉语送气音和日语鼻化音的零样本分类准确率显著提升
  • 不同语音特征需适配不同时间聚合方式,效果更优

近期语音转国际音标(Speech-to-IPA)的发音基础模型依赖图素到音素(G2P)标签,但音素标签未必符合语音学真实。我们评估了在汉语送气音和日语鼻化音上的零样本发音分类任务。一个仅在排除这两种语言的G2P数据上训练的模型,在这两项任务上表现不佳,表明仅靠多语言离散IPA标记不足以应对未见场景。为此,我们提出一种基于每帧提取的连续发音特征(AF)向量的分类方法。该方法优于基于离散标记的方法,尤其在罕见音素上表现突出。进一步发现,针对目标语音差异选择最优的时间聚合方式至关重要:单帧分类对送气音最佳,而段落级分类显著提升鼻化音识别效果。

原文摘要 · Abstract (English)

Recent Phonetic Foundation Models (PFMs) for Speech-to-IPA transcription rely on Grapheme-to-Phoneme (G2P) labels, but the phoneme labels are not necessarily phonetically faithful. To investigate this issue, we evaluate zero-shot phonetic classification on Chinese aspiration and Japanese moraic nasals. A PFM trained on G2P-labeled data excluding these two languages yields poor accuracy on both tasks, showing that multilingual coverage with discrete IPA tokens is not sufficient for unseen settings. To overcome this limitation, we propose a classification method based on continuous Articulatory Feature (AF) vectors extracted from each frame. This AF-based approach outperforms discrete token-based methods, particularly for rare phones. We further show that it is crucial to adopt the optimal temporal aggregation of AF vectors for the target distinction: single-frame classification is best for aspiration, while segmental classification substantially improves nasal classification.

语音识别发音特征零样本学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。