用大模型实时修复语音、智能优先处理紧急呼叫,提升救援效率
Efficient VoIP Communications through LLM-based Real-Time Speech Reconstruction and Call Prioritization for Emergency Services
- 用LLM结合RAG技术重建丢失语音,补全上下文信息
- 在真实场景中实现高精度转写,BLEU/ROUGE评分优异
- 适合应急中心、智能调度系统,缓解人力短缺问题
紧急通信系统常因VoIP中的丢包、带宽限制、信号差、延迟和抖动而中断,导致实时服务质量下降。受困者因恐慌、言语障碍或背景噪声,难以清晰传达关键信息,进一步影响接线员对情况的准确判断。应急中心人手不足也加剧了响应与协调的延迟。本文提出利用大语言模型(LLMs)解决上述挑战:通过重建不完整语音、填补上下文空白,并根据严重程度对来电进行优先排序。系统融合实时转录与检索增强生成(RAG),采用Twilio和AssemblyAI API实现无缝部署。评估显示模型具备高精度,且在BLEU和ROUGE指标上表现良好,与实际需求高度契合,证明其在优化应急响应流程、有效识别关键案件方面的潜力。
原文摘要 · Abstract (English)
Emergency communication systems face disruptions due to packet loss, bandwidth constraints, poor signal quality, delays, and jitter in VoIP systems, leading to degraded real-time service quality. Victims in distress often struggle to convey critical information due to panic, speech disorders, and background noise, further complicating dispatchers' ability to assess situations accurately. Staffing shortages in emergency centers exacerbate delays in coordination and assistance. This paper proposes leveraging Large Language Models (LLMs) to address these challenges by reconstructing incomplete speech, filling contextual gaps, and prioritizing calls based on severity. The system integrates real-time transcription with Retrieval-Augmented Generation (RAG) to generate contextual responses, using Twilio and AssemblyAI APIs for seamless implementation. Evaluation shows high precision, favorable BLEU and ROUGE scores, and alignment with real-world needs, demonstrating the model's potential to optimize emergency response workflows and prioritize critical cases effectively.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。