用音视频对比学习提升发音特征分类准确率
Audio-Vision Contrastive Learning for Phonological Class Recognition
- 融合实时MRI与语音信号,通过对比学习实现多模态融合
- 在USC-TIMIT数据集上达0.81平均F1分数,比单模态提升0.23
- 适合语音病理分析、个性化康复等临床语音技术研究
准确分类发音特征对理解人类言语产生及发展稳健语音技术至关重要,尤其在临床场景中,可提升疾病诊断精度与个性化康复效果。本文提出一种多模态深度学习框架,结合实时磁共振成像(rtMRI)与语音信号,分类三种关键发音维度:发音方式、发音部位和声带振动。在由这些维度导出的15个发音类别上进行评估,采用四种音频/视觉配置:单模态rtMRI、单模态语音信号、多模态中间融合,以及基于对比学习的音视频融合。在USC-TIMIT数据集上的实验结果表明,该对比学习方法达到最优性能,平均F1分数为0.81,相比单模态基线绝对提升0.23。结果验证了对比表示学习在多模态发音分析中的有效性。代码与处理后的数据集将公开于https://github.com/DaE-plz/AC_Contrastive_Phonology,以支持后续研究。
原文摘要 · Abstract (English)
Accurate classification of articulatory-phonological features plays a vital role in understanding human speech production and developing robust speech technologies, particularly in clinical contexts where targeted phonemic analysis and therapy can improve disease diagnosis accuracy and personalized rehabilitation. In this work, we propose a multimodal deep learning framework that combines real-time magnetic resonance imaging (rtMRI) and speech signals to classify three key articulatory dimensions: manner of articulation, place of articulation, and voicing. We perform classification on 15 phonological classes derived from the aforementioned articulatory dimensions and evaluate the system with four audio/vision configurations: unimodal rtMRI, unimodal audio signals, multimodal middle fusion, and contrastive learning-based audio-vision fusion. Experimental results on the USC-TIMIT dataset show that our contrastive learning-based approach achieves state-of-the-art performance, with an average F1-score of 0.81, representing an absolute increase of 0.23 over the unimodal baseline. The results confirm the effectiveness of contrastive representation learning for multimodal articulatory analysis. Our code and processed dataset will be made publicly available at https://github.com/DaE-plz/AC_Contrastive_Phonology to support future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。