微调Whisper模型识别三类语音重音,提升包容性语音识别能力
Fine-Tuning Whisper for Inclusive Prosodic Stress Analysis
- 微调Whisper大模型识别语调、词汇和对比重音
- 三类重音识别接近人类水平,性别与神经类型分类精度近满分
- 适合关注语音识别公平性和多样性研究的学者
语调在语音理解中起关键作用,影响人类认知与自动语音识别(ASR)系统。尽管重要,重音分析仍因效率问题而研究不足。本研究探索微调OpenAI Whisper large-v2 ASR模型,以识别语篇、词汇和对比重音。基于66名母语英语者(含男性、女性、神经正常及神经多样性个体)的数据集,评估模型在不同重音模式上的泛化能力,并通过短语音样本分类说话人神经类型与性别。结果表明,三种重音识别的ASR性能接近人类水平,性别与神经类型分类精度接近完美。该工作提升了语调感知型ASR能力,为多样化人群提供更公平、鲁棒的转录技术。
原文摘要 · Abstract (English)
Prosody plays a crucial role in speech perception, influencing both human understanding and automatic speech recognition (ASR) systems. Despite its importance, prosodic stress remains under-studied due to the challenge of efficiently analyzing it. This study explores fine-tuning OpenAI's Whisper large-v2 ASR model to recognize phrasal, lexical, and contrastive stress in speech. Using a dataset of 66 native English speakers, including male, female, neurotypical, and neurodivergent individuals, we assess the model's ability to generalize stress patterns and classify speakers by neurotype and gender based on brief speech samples. Our results highlight near-human accuracy in ASR performance across all three stress types and near-perfect precision in classifying gender and neurotype. By improving prosody-aware ASR, this work contributes to equitable and robust transcription technologies for diverse populations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。