将音素识别分解为发音特征,提升异常语音识别准确率与可解释性。
Multi-Task Learning for Non-Canonical Phoneme Recognition via Articulatory Feature Decomposition

- 将音素拆解为发音部位、方式、清浊等特征,分任务学习
- 在L2-ARCTIC数据集上比基线模型提升显著,错误模式可解释
- 适合研究语音障碍、跨口音识别的学者与工程师
病理性和更广泛的非标准语音因发音系统性偏离标准发音且标注临床语音数据稀缺,给自动音素识别带来巨大挑战。现有系统通常在标准语音上训练,将音素视为原子类别标签,难以检测语音障碍或口音中常见的结构化发音错误。本文提出一种语言学结构化的非标准音素识别方法,将音素预测分解为发音部位、方式、清浊等发音特征维度。采用层次化多任务学习架构,各特征头学习特征级表征,并通过交叉注意力融合模块生成音素预测。为应对病理语音标签稀缺与噪声问题,结合动量伪标签(MPL)进行半监督学习,并设计级联训练策略:逐步引入发音特征任务,分阶段解冻预训练语音编码器。在L2-ARCTIC数据集(作为病理语音变异的代理)上的实验表明,该方法相比强基线模型显著提升音素识别性能,同时产生与音系特征结构一致的可解释错误模式。结果表明,发音特征监督是提升非标准语音中鲁棒且可解释音素识别的有前景策略,值得在未来临床诊断病理语音数据集上进一步验证。
原文摘要 · Abstract (English)
Pathological and more broadly non-canonical speech present significant challenges for automatic phoneme recognition due to systematic deviations from canonical pronunciation and limited availability of labeled clinical speech data. Existing phoneme recognition systems are typically trained on canonical speech and treat phonemes as atomic categorical labels, limiting their ability to detect structured articulatory errors common in speech disorders and accents. In this work, we introduce a linguistically structured approach to non-canonical phoneme recognition that decomposes phoneme prediction into articulatory feature dimensions such as manner, place, and voicing. We implement this formulation using a hierarchical multi-task learning architecture in which task-specific articulatory feature heads learn feature-level representations that are subsequently integrated through a cross-attention-based fusion module to produce phoneme predictions. To address the scarcity and noise of pathological speech labels, we combine this framework with semi-supervised learning via Momentum Pseudo-Labeling (MPL) and propose a cascaded training strategy that progressively introduces articulatory feature tasks while employing staged unfreezing of a pretrained speech encoder. Experiments on L2-ARCTIC, used as a proxy for pathological speech variation, show that the proposed approach achieves substantial improvements in phoneme recognition performance compared to strong baseline architectures, while yielding interpretable error patterns aligned with phonological feature structure. These results suggest that articulatory feature supervision is a promising strategy for robust and interpretable phoneme recognition in non-canonical speech, and motivate future validation on clinically diagnosed pathological speech datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。