SpeechAgent让有言语障碍者在手机上实时沟通无障碍
SpeechAgent: An End-to-End Mobile Infrastructure for Speech Impairment Assistance
- 融合大模型与语音处理,按障碍类型动态适配支持
- 移动端实时处理延迟极低,准确率和音质均高
- 专为真实患者数据设计,适合日常辅助沟通场景
语言是人类交流的核心,但数百万患有构音障碍、口吃和失语症的人常因此产生社交孤立与参与度下降。尽管自动语音识别(ASR)和文本转语音(TTS)技术近年取得进展,面向言语障碍者的可访问网络与移动基础设施仍十分有限,阻碍了这些技术在日常交流中的实际应用。为此,我们提出SpeechAgent——一个专为言语障碍人群设计的端到端移动端辅助系统。该系统结合大语言模型(LLM)驱动的推理与先进语音处理模块,可根据不同障碍类型提供自适应支持。为确保真实可用性,我们开发了结构化部署流程,实现移动端与边缘设备上的实时语音处理,在保持高准确率与语音质量的同时,达到几乎无感知的延迟。在真实患者语音数据集上的评估及边缘设备延迟分析表明,SpeechAgent在个性化日常沟通中具备高效且用户友好的表现,验证了其可行性。
原文摘要 · Abstract (English)
Speech is essential for human communication, yet millions of people face impairments such as dysarthria, stuttering, and aphasia conditions that often lead to social isolation and reduced participation. Despite recent progress in automatic speech recognition (ASR) and text-to-speech (TTS) technologies, accessible web and mobile infrastructures for users with impaired speech remain limited, hindering the practical adoption of these advances in daily communication. To bridge this gap, we present SpeechAgent, a mobile SpeechAgent designed to facilitate people with speech impairments in everyday communication. The system integrates large language model (LLM)- driven reasoning with advanced speech processing modules, providing adaptive support tailored to diverse impairment types. To ensure real-world practicality, we develop a structured deployment pipeline that enables real-time speech processing on mobile and edge devices, achieving imperceptible latency while maintaining high accuracy and speech quality. Evaluation on real-world impaired speech datasets and edge-device latency profiling confirms that SpeechAgent delivers both effective and user-friendly performance, demonstrating its feasibility for personalized, day-to-day assistive communication.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。