用大模型提升失语者实时语音转换,让表达更清晰自然
Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture

- 三阶段架构:识别-理解-合成,结合大模型增强语音重建
- 在中文失语数据集上显著提升可懂度,保持语义连贯性
- 适合轻中度失语者,为专业场合沟通提供新可能
失语患者在会议、演讲等专业场景中面临严重沟通障碍,现有辅助通信系统因延迟高、语音不自然而难以满足需求。本文提出 Re-Sonance,一种基于 Whisper ASR、Qwen LLM 和 CosyVoice TTS 的三阶段级联语音驱动辅助系统,旨在实现专业场景下的实时高效沟通。在包含轻中度失语者的中文语音数据集上,主观与客观评估均表明,该系统在保持语义一致性的前提下显著提升了语音可懂度,实现了接近实时的响应性能。尽管重度失语者效果仍有限,但研究验证了大模型在语音辅助系统中的潜力,为构建更高效、易用的沟通技术提供了新路径。
原文摘要 · Abstract (English)
Individuals with dysarthria face significant challenges in professional speaking scenarios such as conferences, presentations, and meetings, where real-time communication is crucial. While existing Augmentative and Alternative Communication (AAC) systems provide basic support, they often fail to meet the demands of professional speaking environments due to high latency and unnatural speech patterns. This paper presents Re-Sonance, a novel LLM-enhanced speech-driven AAC system designed for real-time professional speaking scenarios. By integrating Whisper ASR, Qwen LLM, and CosyVoice TTS, Re-Sonance achieves improved speech intelligibility and naturalness while maintaining real-time performance. Both subjective and objective evaluations using a Mandarin dysarthric speech dataset demonstrate that our speech reconstruction approach significantly improved intelligibility while preserving semantic coherence for speakers with mild to moderate dysarthria. Although performance remains limited for severe dysarthria cases, our findings validate the potential of LLM-based methods for enhancing speech-driven AAC systems, paving the way for more effective and accessible communication technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。