用大模型提升失语者语音识别与情绪分析,让沟通更顺畅。
Empowering Dysarthric Speech: Leveraging Advanced LLMs for Accurate Speech Correction and Multimodal Emotion Analysis
- 用Whisper转录+大模型纠错,重建失语者原意
- 在混合数据集上实现7种情绪精准识别
- 适合康复医疗与无障碍技术研究者
失语症是由神经系统损伤导致的运动性言语障碍,影响全球数百万人,表现为发音含糊、缓慢或难以理解。本文提出一种新方法,通过先进大语言模型实现失语语音的准确纠正与多模态情绪分析。首先使用OpenAI Whisper将失语语音转为文本,再通过微调的开源模型及GPT-4o、LLaMA 3.1 70B、Mistral 8x7B等模型在Groq AI加速器上进行句子预测。数据集融合TORGO与谷歌语音数据,并人工标注情感标签。系统可识别幸福、悲伤、中性、惊讶、愤怒、恐惧共7类情绪,同时高精度重构被扭曲的原句。该框架显著提升了失语语音的识别与理解能力。
原文摘要 · Abstract (English)
Dysarthria is a motor speech disorder caused by neurological damage that affects the muscles used for speech production, leading to slurred, slow, or difficult-to-understand speech. It affects millions of individuals worldwide, including those with conditions such as stroke, traumatic brain injury, cerebral palsy, Parkinsons disease, and multiple sclerosis. Dysarthria presents a major communication barrier, impacting quality of life and social interaction. This paper introduces a novel approach to recognizing and translating dysarthric speech, empowering individuals with this condition to communicate more effectively. We leverage advanced large language models for accurate speech correction and multimodal emotion analysis. Dysarthric speech is first converted to text using OpenAI Whisper model, followed by sentence prediction using fine-tuned open-source models and benchmark models like GPT-4.o, LLaMA 3.1 70B and Mistral 8x7B on Groq AI accelerators. The dataset used combines the TORGO dataset with Google speech data, manually labeled for emotional context. Our framework identifies emotions such as happiness, sadness, neutrality, surprise, anger, and fear, while reconstructing intended sentences from distorted speech with high accuracy. This approach demonstrates significant advancements in the recognition and interpretation of dysarthric speech.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。