跨语言语音评估框架,精准捕捉发音异常模式。
Multilingual Dysarthric Speech Assessment Using Universal Phone Recognition and Language-Specific Phonemic Contrast Modeling
- 用对比声学特征实现跨语言音素映射与对齐
- 三种新指标:PER、PFER、PhonCov,分别反映不同层面错误
- 适用于英语、西班牙语、意大利语、泰米尔语,临床相关性强
神经性构音障碍患病率上升,推动跨语言自动化发音清晰度评估需求。现有方法多局限于单一语言,或无法捕捉语言特异性影响因素。本文提出一种多语言音素生成评估框架,结合通用音素识别与语言特异性音素解释,利用对比音位特征距离实现音素到音位的映射与序列对齐。该框架生成三项指标:音素错误率(PER)、音位特征错误率(PFER)及新提出的无对齐度量——音素覆盖率(PhonCov)。在英语、西班牙语、意大利语和泰米尔语上的分析表明,PER 受映射与对齐共同提升,PFER 仅受益于对齐,PhonCov 仅受益于映射。进一步分析显示,该框架捕捉到与已知构音障碍规律一致的清晰度退化模式,具有临床意义。
原文摘要 · Abstract (English)
The growing prevalence of neurological disorders associated with dysarthria motivates the need for automated intelligibility assessment methods that are applicalbe across languages. However, most existing approaches are either limited to a single language or fail to capture language-specific factors shaping intelligibility. We present a multilingual phoneme-production assessment framework that integrates universal phone recognition with language-specific phoneme interpretation using contrastive phonological feature distances for phone-to-phoneme mapping and sequence alignment. The framework yields three metrics: phoneme error rate (PER), phonological feature error rate (PFER), and a newly proposed alignment-free measure, phoneme coverage (PhonCov). Analysis on English, Spanish, Italian, and Tamil show that PER benefits from the combination of mapping and alignment, PFER from alignment alone, and PhonCov from mapping. Further analyses demonstrate that the proposed framework captures clinically meaningful patterns of intelligibility degradation consistent with established observations of dysarthric speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。