Harf-Speech实现阿拉伯语音素级发音评估,临床相关性强。
Harf-Speech: A Clinically Aligned Framework for Arabic Phoneme-Level Speech Assessment

- 融合音素转换器与语音转音素模型,结合编辑距离等度量方法。
- 最佳模型在音素错误率上达8.92%,与专家评分相关性0.791。
- 适合语言治疗、阿拉伯语教学及临床发音评估研究者使用。
自动化音素级发音评估对大规模言语治疗和语言学习至关重要,但针对阿拉伯语的验证工具仍十分稀缺。本文提出 Harf-Speech,一个模块化系统,可在临床量表上实现阿拉伯语发音的音素级评分。该系统结合标准阿拉伯语(MSA)音素转换器、微调的语音到音素模型、Levenshtein 对齐算法,以及基于最长公共子序列与编辑距离的混合评分机制。我们在阿拉伯语音素数据上微调了三种 ASR 架构,并与零样本多模态模型进行对比;其中表现最佳的 OmniASR-CTC-1B-v2 模型达到 8.92% 的音素错误率。三位认证的语言病理学家独立评估了 40 个发音片段作为临床验证。Harf-Speech 与平均专家评分的皮尔逊相关系数为 0.791,组内相关系数 ICC(2,1) 为 0.659,优于现有端到端评估框架。结果表明,Harf-Speech 能生成与临床一致且可解释的评分,接近专家间一致性水平。
原文摘要 · Abstract (English)
Automated phoneme-level pronunciation assessment is vital for scalable speech therapy and language learning, yet validated tools for Arabic remain scarce. We present Harf-Speech, a modular system scoring Arabic pronunciation at the phoneme level on a clinical scale. It combines an MSA phonetizer, a fine-tuned speech-to-phoneme model, Levenshtein alignment, and a blended scorer using longest common subsequence and edit-distance metrics. We fine-tune three ASR architectures on Arabic phoneme data and benchmark them with zero-shot multimodal models; the best, OmniASR-CTC-1B-v2, achieves 8.92% phoneme error rate. Three certified speech-language pathologists independently scored 40 utterances for clinical validation. Harf-Speech attains a Pearson correlation of 0.791 and ICC(2,1) of 0.659 with mean expert scores, outperforming existing end-to-end assessment frameworks. These results show Harf-Speech yields clinically aligned, interpretable scores comparable to inter-rater expert agreement.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。