arXiv:2605.25596cs.CL2026-05中稿 · Interspeech 2026被引 1

用自监督模型直接识别多语言语音的22维音系特征,准确率超传统方法。

Multilingual Phonological Feature Recognition with Self-Supervised Speech Models

论文配图:Multilingual Phonological Feature Recognition with Self-Supervised Speech Models
图 1 · 摘自论文原文
  • 基于自监督语音模型,直接预测每帧22维音系特征向量。
  • 跨语言测试中平均宏F1达88.9%,比基线提升8.6点。
  • 在未见语言上性能显著提升,适合多语言语音分析研究者。

音系特征提供了语言通用且符合语言学原理的语音表征。我们提出PhonoQ-2.0,一个基于自监督语音模型的多语言帧级音系特征识别系统。该系统直接预测每帧包含发音方式、元音质量、发音部位和清浊的22维特征向量,而非从音素输出推导。为确保音系一致性,引入一种方式条件门控机制,激活有效特征组。在多个语言和语料上评估,PhonoQ-2.0在域内平均宏F1达到91.3%,域外达88.9%。相比强基线CTC音素模型,域内平均提升+8.8 F1,域外+8.6。在未见语言评估中,宏F1从66.9%提升至73.6%(平均+6.7),最高提升达+10.8。

原文摘要 · Abstract (English)

Phonological features provide a language-general and linguistically grounded representation of speech. We present PhonoQ-2.0, a multilingual frame-level phonological feature recognizer built on self-supervised speech models. The system directly predicts a structured 22-dimensional feature vector per frame encoding manner, vowel quality, place, and voicing, instead of deriving features from phoneme outputs. To ensure phonologically coherent predictions, we introduce a manner-conditioned gating mechanism that activates valid feature groups. Evaluated across multiple languages and corpora, PhonoQ-2.0 achieves an average macro-F1 of 91.3% in-domain and 88.9% out-of-domain. Compared to a strong CTC phoneme baseline, it delivers consistent gains of +8.8 F1 in-domain and +8.6 out-of-domain on average. In unseen-language evaluation, PhonoQ-2.0 improves macro-F1 from 66.9% to 73.6% (+6.7 on average), with gains of up to +10.8 points.

音系特征自监督学习多语言语音识别

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。